Skip to content

Model comparison

Claude Opus 5.5 vs Claude Haiku 4.5: cost and context limits

Claude Opus 5.5 lists at 4x the rates of Claude Haiku 4.5, yet a cached coding session costs 3.7x. Context and output limits, and Haiku 4.5's retirement date.

· Prices as of September 26, 2026

  • Claude Opus 5.5

    Anthropic · Released September 22, 2026

    Anthropic's recommended starting model for most work, built for long-running agentic coding. Anthropic says it matches Claude Fable 5.1 on most work at a much lower price.

    Claude Opus 5.5 facts and comparisons
  • Claude Haiku 4.5

    Anthropic · Released October 15, 2025

    Anthropic's fastest and cheapest current model, aimed at latency-sensitive work and at coding sub-agents.

    Claude Haiku 4.5 facts and comparisons

The short answer

Claude Haiku 4.5 lists at a quarter of Claude Opus 5.5's rates, but Opus 5.5's discounted cache hits narrow the example agentic coding session to 3.7x: $4.40 against $1.20. Opus 5.5 is Anthropic's recommended starting model and Claude Code's default, with a 1M context window and 128K of output. Haiku 4.5, capped at 200K of context and 64K of output, is Anthropic's cheapest current model, aimed at latency-sensitive work and coding sub-agents.

Choose Claude Opus 5.5 if

  • You need more than 200K tokens of context for codebase-wide migrations or audits, which Anthropic names as an Opus 5.5 strength.
  • You want Claude Code's default model and Anthropic's recommended starting point for most work.
  • Your sessions resend large contexts, where an Opus 5.5 cache hit costs $0.20 per million, only 2x Haiku 4.5's $0.10.

Choose Claude Haiku 4.5 if

  • You run sub-agents or high-volume tasks that fit in 200K tokens of context and 64K of output.
  • You send many short uncached requests, where Opus 5.5 costs the full 4x.
  • You want to cap thinking with a fixed token budget rather than adaptive thinking.

Side by side

Specs and prices

FactClaude Opus 5.5Claude Haiku 4.5
MakerAnthropicAnthropic
API model idclaude-opus-5-5claude-haiku-4-5
ReleasedSeptember 22, 2026October 15, 2025
StatusCurrentCurrent
Context window1M tokens200K tokens
Max output128K tokens64K tokens
Open weightsNoNo
Input, per 1M tokens$4$1
Cache hit, per 1M$0.20$0.10
Cache write, per 1M$5 (5-minute), $8 (1-hour)$1.25 (5-minute), $2 (1-hour)
Output, per 1M tokens$20$5
Runs inClaude Code, Cursor, OpenCode, OpenRouter, and GitHub CopilotClaude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker on September 26, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Claude Opus 5.5: The full 1M context window is billed at standard rates. Fast mode, a research preview on the Claude API, costs $8 input and $40 output per million tokens.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadClaude Opus 5.5Claude Haiku 4.5
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$4.40$1.20
Large one-off review, 150K input with no cache hits, 10K output$0.80$0.20
Output-heavy generation, 30K input, 80K output$1.72$0.43
A month of sessions, 110 sessions: 5 a day, 22 working days$484.00$132.00
Where the session’s cost goes
Cache writes$2.60$0.65
Cache reads$0.40$0.20
Uncached input$0.40$0.10
Output$1.00$0.25
caching saves on the session with Claude Opus 5.5 (60%)
$6.60
caching saves on the session with Claude Haiku 4.5 (56%)
$1.55

Why a 4x rate gap becomes 3.7x on a session

Claude Opus 5.5 charges $4 per million input tokens and $20 per million output tokens. Claude Haiku 4.5 charges $1 and $5. Cache writes follow the same 4x: $5 and $8 for Opus 5.5's 5-minute and 1-hour writes, $1.25 and $2 for Haiku 4.5's. Uncached work pays that multiple in full, so the large one-off review costs $0.80 against $0.20 and the output-heavy generation $1.72 against $0.43.

Cache hits break the pattern. Anthropic bills an Opus 5.5 hit at 0.05x its input price, half the 0.1x it charges on Haiku 4.5 and most other Claude models. That makes a hit $0.20 on Opus 5.5 and $0.10 on Haiku 4.5, a 2x gap instead of 4x. The 2M cached tokens in the example session cost $0.40 and $0.20.

Everything else in the session still runs at 4x, so the total lands at 3.7x, a $3.20 difference per session. At 110 sessions a month that is $484.00 against $132.00 at API rates. The largest line on both is cache writes, $2.60 on Opus 5.5 and $0.65 on Haiku 4.5, which is 59% and 54% of each session.

Context, output, and tokenizer differences

Opus 5.5 accepts 1M tokens of context at standard rates, five times Haiku 4.5's 200K, and writes up to 128K tokens of output against 64K. Anthropic names long, sprawling jobs such as codebase-wide migrations and audits as a particular Opus 5.5 strength, and those are the tasks most likely to need more than 200K tokens in a single request.

The token counts themselves aren't directly comparable. Haiku 4.5 uses Anthropic's older tokenizer, and Anthropic notes that Claude models from 4.7 on, Opus 5.5 among them, produce about 30% more tokens for the same text. So the same file costs more tokens on Opus 5.5 as well as more per token, and 200K on Haiku 4.5 holds more text than 200K would on Opus 5.5.

Using Opus 5.5 and Haiku 4.5 together

Anthropic positions the two for different jobs. Opus 5.5 is its recommended starting model for most work, and Anthropic says it performs at the level of Claude Fable 5.1 on most work. Haiku 4.5 is aimed at latency-sensitive work and coding sub-agents, and Anthropic says its coding performance is similar to Claude Sonnet 4 at one-third the cost and more than twice the speed.

That split maps onto multi-agent coding: Opus 5.5 running the main conversation, which it does by default in Claude Code, and Haiku 4.5 handling sub-agents in refactors and migrations, a use Anthropic names for it. Thinking works differently on each. Opus 5.5 has adaptive thinking always on at medium effort by default, while Haiku 4.5 uses manual extended thinking with a token budget.

Haiku 4.5 also has a horizon. Anthropic lists its retirement as not sooner than October 15, 2026, and announced Claude Haiku 5.5 on September 22, 2026, so a sub-agent setup built on Haiku 4.5 now should expect a revisit. To see how a mixed setup splits in practice, EveryToken prices each request in your local Claude Code history at Anthropic's rates and totals cost by model.

Prompt caching

How Anthropic bills cached tokens

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

Your own numbers

See what Claude Opus 5.5 and Claude Haiku 4.5 really cost you.

everyaitoken reads your Claude Code, Cursor, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

How much more does Claude Opus 5.5 cost than Claude Haiku 4.5?

4x per token for input, output, and cache writes, and 2x for cache hits. The example agentic session costs $4.40 on Opus 5.5 and $1.20 on Haiku 4.5, a 3.7x gap.

Can Claude Haiku 4.5 use a 1M token context window?

No. Its context window is 200K tokens, with up to 64K of output. Opus 5.5 takes 1M tokens and writes up to 128K.

Which of the two does Claude Code use by default?

Claude Opus 5.5, on Pro, Max, Team, and Enterprise plans and with an Anthropic API key. Haiku 4.5 is available in Claude Code too, and you can select it.

What is replacing Claude Haiku 4.5?

Claude Haiku 5.5 was announced on September 22, 2026. It lists Haiku 4.5's retirement as not sooner than October 15, 2026.

  • Claude Fable 5.1 vs Claude Opus 5.5

    Claude Fable 5.1 lists at 2.5x the rates of Claude Opus 5.5. What that premium buys, why cheap cache hits barely narrow it, and when Anthropic suggests it.

  • Claude Opus 5.5 vs Claude Opus 4.8

    Claude Opus 5.5 lists 20% below Claude Opus 4.8 and bills cache hits at $0.20, not $0.50. What staying costs, and when Anthropic still recommends Opus 4.8.

  • Claude Opus 5.5 vs Claude Opus 5

    Claude Opus 5.5 is 20% cheaper per token than Claude Opus 5, and cache hits cost 60% less. What that means for Claude Code sessions, and what stays the same.

  • Claude Opus 5.5 vs Claude Sonnet 5

    Claude Opus 5.5 costs twice as much per token as Claude Sonnet 5, yet cache hits cost $0.20 on both. What that means for agentic coding sessions.

  • Claude Sonnet 5 vs Claude Haiku 4.5

    Claude Haiku 4.5 costs half as much as Claude Sonnet 5, with a 200K context window and a retirement date ahead. When the cheaper model fits a coding workflow.

  • Claude Haiku 4.5 vs Gemini 3.5 Flash-Lite

    Claude Haiku 4.5 and Gemini 3.5 Flash-Lite both target sub-agent work. Haiku 4.5 costs 3.5x as much per cached session, but only 2x as much per output token.