Skip to content

Model comparison

Claude Sonnet 5 vs GPT-5.6 Terra: the cheaper pick flips

Claude Sonnet 5 and GPT-5.6 Terra share a $2 input rate. Terra costs less on cached sessions, Sonnet 5 on output-heavy work. The numbers, line by line.

· Prices as of September 28, 2026

  • Claude Sonnet 5

    Anthropic · Released June 30, 2026

    Anthropic's balance of speed and intelligence, and a drop-in upgrade for Claude Sonnet 4.6. Anthropic pitches it as close to Claude Opus 4.8 at a lower price.

    Claude Sonnet 5 facts and comparisons
  • GPT-5.6 Terra

    OpenAI · Released July 9, 2026 · Previous generation

    The middle tier of GPT-5.6, balancing intelligence and cost, which roughly corresponds to earlier GPT-5 mini models.

    GPT-5.6 Terra facts and comparisons

The short answer

Claude Sonnet 5 and GPT-5.6 Terra cost almost the same, and which is cheaper depends on the work. GPT-5.6 Terra is $0.20 cheaper on the example agentic coding session, $2.20 against $2.40, while Sonnet 5 is $0.16 cheaper on output-heavy generation because it charges $10 per million output tokens to Terra's $12. Terra is a previous model that Codex suggests replacing with GPT-6 Sol, so for new work the choice mostly comes down to your tools.

Choose Claude Sonnet 5 if

  • Your work is output-heavy, such as generating new files, tests, or long explanations, where Sonnet 5's $10 output rate beats Terra's $12.
  • You call Claude through an API key, where Claude Code's main conversation uses the 5-minute cache and Sonnet 5's writes cost the same as Terra's.
  • You want a current model rather than one Codex already suggests replacing.
  • You are upgrading from Claude Sonnet 4.6, which Anthropic says Sonnet 5 replaces as a drop-in upgrade.

Choose GPT-5.6 Terra if

  • Your sessions write a lot of context to the cache and you would otherwise pay Anthropic's 2x rate for 1-hour writes.
  • You already work in Codex or call gpt-5.6-terra from code, and it stays available in the API.
  • You like OpenAI's pitch that GPT-5.6 infers your underlying goal and intended level of work from context.
  • You were on an earlier GPT-5 mini model, the tier OpenAI says Terra roughly corresponds to.

Side by side

Specs and prices

FactClaude Sonnet 5GPT-5.6 Terra
MakerAnthropicOpenAI
API model idclaude-sonnet-5gpt-5.6-terra
ReleasedJune 30, 2026July 9, 2026
StatusCurrentPrevious generation
Context window1M tokens1.05M tokens
Max output128K tokens128K tokens
Open weightsNoNo
Input, per 1M tokens$2$2
Cache hit, per 1M$0.20$0.20
Cache write, per 1M$2.50 (5-minute), $4 (1-hour)$2.50
Output, per 1M tokens$10$12
Runs inClaude Code, Cursor, OpenCode, OpenRouter, and GitHub CopilotCodex, Cursor, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker (Claude Sonnet 5: September 26, 2026; GPT-5.6 Terra: September 28, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Claude Sonnet 5: The full 1M context window is billed at standard rates. The launch price became the standard price on August 10, 2026, and a planned increase was cancelled. GPT-5.6 Terra: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadClaude Sonnet 5GPT-5.6 Terra
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$2.40$2.20
Large one-off review, 150K input with no cache hits, 10K output$0.40$0.42
Output-heavy generation, 30K input, 80K output$0.86$1.02
A month of sessions, 110 sessions: 5 a day, 22 working days$264.00$242.00
Where the session’s cost goes
Cache writes$1.30$1.00
Cache reads$0.40$0.40
Uncached input$0.20$0.20
Output$0.50$0.60
caching saves on the session with Claude Sonnet 5 (56%)
$3.10
caching saves on the session with GPT-5.6 Terra (61%)
$3.40

Where GPT-5.6 Terra is cheaper, and where Claude Sonnet 5 is

Most of the two rate cards match. Claude Sonnet 5 and GPT-5.6 Terra both charge $2 per million input tokens, $0.20 per million cache hits, and $2.50 per million for a standard cache write. They differ in two places: Terra charges $12 per million output tokens where Sonnet 5 charges $10 per million, 20% more, and Anthropic also sells a 1-hour cache write at $4 per million.

Those two differences pull in opposite directions. On the large one-off review, with 10K tokens of output, Sonnet 5 costs $0.40 and Terra $0.42. On the output-heavy generation, with 80K output tokens, Sonnet 5 costs $0.86 and Terra $1.02, a 16% saving for Sonnet 5. On the cached agentic session Terra wins, $2.20 against $2.40.

What decides the session: cache writes or output

In the session, Sonnet 5 pays $0.30 more for cache writes ($1.30 against $1.00), because half of its 400K written tokens go to the 1-hour cache at 2x input. Terra pays $0.10 more for output ($0.60 against $0.50). Reads and fresh input cost the same. The net is $0.20 in Terra's favor, or $22.00 over 110 sessions a month, $242.00 against $264.00.

That result depends on the cache lifetime. Claude Code writes the main conversation to the 1-hour cache on a Claude subscription and to the 5-minute cache when you use an API key. With only 5-minute writes, Sonnet 5's write cost would match Terra's and its lower output rate would make it the cheaper of the two on this session as well.

Caching saves $3.40 on Terra, 61% of the uncached cost, and $3.10 on Sonnet 5, or 56%. The cache walkthrough on our homepage shows how the same session breaks down line by line.

Terra's status, its 20% price cut, and the limits

OpenAI cut Terra's price by 20% on July 30, 2026, and the rates in this post are the reduced ones. It is now a previous model: Codex suggests moving to GPT-6 Sol, though Terra stays in the API and in Cursor, OpenRouter, OpenCode, and GitHub Copilot. OpenAI calls it the "GPT-5.6 model that balances intelligence and cost."

Anthropic's pitch for Sonnet 5 is agentic work: planning and using tools like browsers and terminals on its own, with performance it describes as close to Claude Opus 4.8. Adaptive thinking is on by default at high effort, and higher effort usually means more output tokens, which is the one rate where Sonnet 5 has the edge.

Limits are close. Sonnet 5 has a 1M context window and Terra 1.05M, and both write up to 128K tokens. Terra's rates rise for any request over 272K input tokens: 2x for input and cache and 1.5x for output, across the whole request.

Prompt caching

How each maker bills cached tokens

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

OpenAI

Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.

From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.

On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.

Source: OpenAI: Prompt caching

Your own numbers

See what Claude Sonnet 5 and GPT-5.6 Terra really cost you.

everyaitoken reads your Claude Code, Cursor, OpenCode, OpenRouter, and Codex history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is GPT-5.6 Terra cheaper than Claude Sonnet 5?

On cached agentic sessions, slightly: $2.20 against $2.40 in the example. On uncached or output-heavy work Sonnet 5 costs less, $0.86 against $1.02 for the output-heavy generation. Input and cache-hit rates are identical.

Will GPT-5.6 Terra go away?

No retirement date is listed. Codex suggests moving to GPT-6 Sol, and Terra stays available in the API.

Why does the cache lifetime change which model is cheaper?

Anthropic charges 1.25x input for a 5-minute cache write and 2x for a 1-hour write, while OpenAI charges 1.25x for every write from GPT-5.6 on. When Claude uses the 1-hour cache, Sonnet 5's writes cost more than Terra's, and that outweighs its cheaper output on a cache-heavy session.

How can I compare them on my own sessions?

EveryToken reads Claude Code, Codex, and Cursor history on your Mac, prices each request at API rates, and shows what caching saved or cost per model. That tells you whether your work leans toward cache writes or output.

  • Claude Fable 5.1 vs Claude Sonnet 5

    Claude Fable 5.1 lists at 5x the rates of Claude Sonnet 5, but cache hits cost $0.25 against $0.20. Why sessions cost 4.4x, and who Anthropic aims Fable 5.1 at.

  • Claude Haiku 4.5 vs GPT-5.6 Terra

    Claude Haiku 4.5 costs about half of GPT-5.6 Terra per token, $1.20 against $2.20 per coding session. Terra offers a 1.05M context window and 128K output.

  • Claude Opus 5.5 vs Claude Sonnet 5

    Claude Opus 5.5 costs twice as much per token as Claude Sonnet 5, yet cache hits cost $0.20 on both. What that means for agentic coding sessions.

  • Claude Sonnet 5 vs Claude Haiku 4.5

    Claude Haiku 4.5 costs half as much as Claude Sonnet 5, with a 200K context window and a retirement date ahead. When the cheaper model fits a coding workflow.

  • Claude Sonnet 5 vs Claude Sonnet 4.6

    Claude Sonnet 5 costs 33% less per token than Claude Sonnet 4.6, but its tokenizer counts up to 1.35x as many tokens. What the upgrade saves in practice.

  • Claude Sonnet 5 vs GPT-5.5

    GPT-5.5 costs 2.5x Claude Sonnet 5 for input and 3x for output, and it leaves ChatGPT and Codex sign-in on October 14, 2026. What that means for coding.