Skip to content

Model comparison

Claude Haiku 4.5 vs GPT-5.6 Terra: small model, middle price

Claude Haiku 4.5 costs about half of GPT-5.6 Terra per token, $1.20 against $2.20 per coding session. Terra offers a 1.05M context window and 128K output.

· Prices as of September 28, 2026

  • Claude Haiku 4.5

    Anthropic · Released October 15, 2025

    Anthropic's fastest and cheapest current model, aimed at latency-sensitive work and at coding sub-agents.

    Claude Haiku 4.5 facts and comparisons
  • GPT-5.6 Terra

    OpenAI · Released July 9, 2026 · Previous generation

    The middle tier of GPT-5.6, balancing intelligence and cost, which roughly corresponds to earlier GPT-5 mini models.

    GPT-5.6 Terra facts and comparisons

The short answer

Claude Haiku 4.5 is the cheaper model, at $1 input and $5 output per million against $2 and $12 for GPT-5.6 Terra, and the example agentic coding session costs $1.20 against $2.20. GPT-5.6 Terra is OpenAI's middle GPT-5.6 tier with a 1.05M context window and 128K of output, against Haiku 4.5's 200K and 64K. Pick Haiku 4.5 for low-cost work that fits its window, and Terra when a task needs the extra room.

Choose Claude Haiku 4.5 if

  • You want to spend less: Haiku 4.5's input rate is half of Terra's and its output rate is $5 against $12.
  • Your work is output-heavy, where the gap is widest: $0.43 against $1.02 on the output-heavy generation.
  • You run sub-agents in Claude Code, which Anthropic names as a core use for Haiku 4.5.

Choose GPT-5.6 Terra if

  • Your tasks need more than 200K tokens of context or more than 64K tokens of output.
  • You work in Codex and want a GPT-5.6 model, although Codex now suggests GPT-6 Sol.
  • You want a cached prefix that stays reusable for at least 30 minutes after its last use, at a single 1.25x write price.
  • You like OpenAI's claim that GPT-5.6 infers the user's underlying goal and intended level of work from context.

Side by side

Specs and prices

FactClaude Haiku 4.5GPT-5.6 Terra
MakerAnthropicOpenAI
API model idclaude-haiku-4-5gpt-5.6-terra
ReleasedOctober 15, 2025July 9, 2026
StatusCurrentPrevious generation
Context window200K tokens1.05M tokens
Max output64K tokens128K tokens
Open weightsNoNo
Input, per 1M tokens$1$2
Cache hit, per 1M$0.10$0.20
Cache write, per 1M$1.25 (5-minute), $2 (1-hour)$2.50
Output, per 1M tokens$5$12
Runs inClaude Code, Cursor, OpenCode, OpenRouter, and GitHub CopilotCodex, Cursor, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker (Claude Haiku 4.5: September 26, 2026; GPT-5.6 Terra: September 28, 2026). Batch and priority tiers, taxes, and subscription plans are not included. GPT-5.6 Terra: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadClaude Haiku 4.5GPT-5.6 Terra
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$1.20$2.20
Large one-off review, 150K input with no cache hits, 10K output$0.20$0.42
Output-heavy generation, 30K input, 80K output$0.43$1.02
A month of sessions, 110 sessions: 5 a day, 22 working days$132.00$242.00
Where the session’s cost goes
Cache writes$0.65$1.00
Cache reads$0.20$0.40
Uncached input$0.10$0.20
Output$0.25$0.60
caching saves on the session with Claude Haiku 4.5 (56%)
$1.55
caching saves on the session with GPT-5.6 Terra (61%)
$3.40

Why this is a cross-tier comparison

These two sit at different levels of their makers' lineups. Claude Haiku 4.5 is Anthropic's lowest-priced current model, aimed at latency-sensitive work and coding sub-agents. GPT-5.6 Terra is the middle tier of GPT-5.6, which OpenAI describes as balancing intelligence and cost and which roughly corresponds to earlier GPT-5 mini models.

The rates reflect that. Terra charges $2 per million input tokens, $0.20 per million cache hits, $2.50 per million cache writes, and $12 per million output tokens, after a 20% price cut on July 30, 2026. Haiku 4.5 charges exactly half on input, hits, and 5-minute writes, and its $5 output rate is 58% below Terra's.

How the gap changes with the workload

Output is where the gap is widest. The output-heavy generation costs $1.02 on Terra and $0.43 on Haiku 4.5, 2.4x. The large one-off review, which is mostly input, costs $0.42 against $0.20, 2.1x.

The cached agentic session narrows it to 1.8x, $2.20 against $1.20. Anthropic's 1-hour cache write costs 2x input, $2 per million on Haiku 4.5, and Terra's single write rate of $2.50 is only 1.25x that. So on the half of the session's writes that go to the 1-hour cache, Haiku 4.5 loses most of its price advantage. Writes cost $0.65 against $1.00, a $0.35 gap, the same size as the output gap of $0.35.

Over 110 sessions a month the totals are $132.00 and $242.00, a $110.00 difference in API-equivalent cost. Caching saves $3.40 per session on Terra and $1.55 on Haiku 4.5.

Limits, status, and tools

Terra has 5.3x the context and twice the output: 1.05M against 200K, and 128K against 64K. Past 272K input tokens, Terra's whole request is billed at 2x for input and cache and 1.5x for output, but that is still a size Haiku 4.5 cannot accept at all.

Both models have successors in view. Codex suggests moving from Terra to GPT-6 Sol, and Terra stays in the API. On Anthropic's side, Haiku 4.5 retires no sooner than October 15, 2026, and Anthropic announced Claude Haiku 5.5 on September 22, 2026.

Each is offered in Cursor, OpenRouter, OpenCode, and GitHub Copilot. Haiku 4.5 is in Claude Code, where it uses manual extended thinking with a token budget. Terra is in Codex. Haiku 4.5 uses Anthropic's older tokenizer, and OpenAI's counts text differently, so identical prompts will not produce identical token totals.

Prompt caching

How each maker bills cached tokens

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

OpenAI

Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.

From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.

On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.

Source: OpenAI: Prompt caching

Your own numbers

See what Claude Haiku 4.5 and GPT-5.6 Terra really cost you.

everyaitoken reads your Claude Code, Cursor, OpenCode, OpenRouter, and Codex history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is Claude Haiku 4.5 cheaper than GPT-5.6 Terra?

Yes, on every rate and every workload here. The example agentic session costs $1.20 against $2.20, and a month of 110 sessions $132.00 against $242.00.

Did GPT-5.6 Terra get a price cut?

OpenAI cut its price by 20% on July 30, 2026. The rates in this post are the reduced ones: $2 input and $12 output per million tokens.

What if my task doesn't fit in 200K tokens?

Claude Haiku 4.5 cannot take it, and GPT-5.6 Terra can, up to its 1.05M window. Requests over 272K input tokens on Terra are billed at its long-context rates.

How can I see which model my team spends more on?

EveryToken reads local history from Claude Code, Codex, and Cursor, prices each request at API rates, and shows the cost and cache savings per model. It is a $9 one-time app for macOS.

  • Claude Haiku 4.5 vs GPT-5.6 Luna

    After an 80% price cut, GPT-5.6 Luna costs a fifth of Claude Haiku 4.5 for input. A month of cached coding sessions: $24.20 against $132.00.

  • Claude Haiku 4.5 vs GPT-6 Luna

    GPT-6 Luna lists a tenth of Claude Haiku 4.5's rates and a 1.05M context window against 200K. What each costs for sub-agents and cached coding sessions.

  • Claude Opus 5.5 vs Claude Haiku 4.5

    Claude Opus 5.5 lists at 4x the rates of Claude Haiku 4.5, yet a cached coding session costs 3.7x. Context and output limits, and Haiku 4.5's retirement date.

  • Claude Sonnet 5 vs Claude Haiku 4.5

    Claude Haiku 4.5 costs half as much as Claude Sonnet 5, with a 200K context window and a retirement date ahead. When the cheaper model fits a coding workflow.

  • Claude Sonnet 5 vs GPT-5.6 Terra

    Claude Sonnet 5 and GPT-5.6 Terra share a $2 input rate. Terra costs less on cached sessions, Sonnet 5 on output-heavy work. The numbers, line by line.

  • GPT-5.6 Sol vs GPT-5.6 Terra

    GPT-5.6 Sol costs about twice GPT-5.6 Terra, on promotional rates. How Terra's output price narrows the gap and why Codex points both to GPT-6 Sol.