Model comparison
Claude Opus 5.5 vs Claude Haiku 4.5: cost and context limits
Claude Opus 5.5 lists at 4x the rates of Claude Haiku 4.5, yet a cached coding session costs 3.7x. Context and output limits, and Haiku 4.5's retirement date.
· Prices as of September 26, 2026
Claude Opus 5.5
Anthropic · Released September 22, 2026
Anthropic's recommended starting model for most work, built for long-running agentic coding. Anthropic says it matches Claude Fable 5.1 on most work at a much lower price.
Claude Opus 5.5 facts and comparisonsClaude Haiku 4.5
Anthropic · Released October 15, 2025
Anthropic's fastest and cheapest current model, aimed at latency-sensitive work and at coding sub-agents.
Claude Haiku 4.5 facts and comparisons
The short answer
Claude Haiku 4.5 lists at a quarter of Claude Opus 5.5's rates, but Opus 5.5's discounted cache hits narrow the example agentic coding session to 3.7x: $4.40 against $1.20. Opus 5.5 is Anthropic's recommended starting model and Claude Code's default, with a 1M context window and 128K of output. Haiku 4.5, capped at 200K of context and 64K of output, is Anthropic's cheapest current model, aimed at latency-sensitive work and coding sub-agents.
Choose Claude Opus 5.5 if
- You need more than 200K tokens of context for codebase-wide migrations or audits, which Anthropic names as an Opus 5.5 strength.
- You want Claude Code's default model and Anthropic's recommended starting point for most work.
- Your sessions resend large contexts, where an Opus 5.5 cache hit costs $0.20 per million, only 2x Haiku 4.5's $0.10.
Choose Claude Haiku 4.5 if
- You run sub-agents or high-volume tasks that fit in 200K tokens of context and 64K of output.
- You send many short uncached requests, where Opus 5.5 costs the full 4x.
- You want to cap thinking with a fixed token budget rather than adaptive thinking.
Side by side
Specs and prices
| Fact | Claude Opus 5.5 | Claude Haiku 4.5 |
|---|---|---|
| Maker | Anthropic | Anthropic |
| API model id | claude-opus-5-5 | claude-haiku-4-5 |
| Released | September 22, 2026 | October 15, 2025 |
| Status | Current | Current |
| Context window | 1M tokens | 200K tokens |
| Max output | 128K tokens | 64K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $4 | $1 |
| Cache hit, per 1M | $0.20 | $0.10 |
| Cache write, per 1M | $5 (5-minute), $8 (1-hour) | $1.25 (5-minute), $2 (1-hour) |
| Output, per 1M tokens | $20 | $5 |
| Runs in | Claude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot | Claude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker on September 26, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Claude Opus 5.5: The full 1M context window is billed at standard rates. Fast mode, a research preview on the Claude API, costs $8 input and $40 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Claude Opus 5.5 | Claude Haiku 4.5 |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $4.40 | $1.20 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.80 | $0.20 |
| Output-heavy generation, 30K input, 80K output | $1.72 | $0.43 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $484.00 | $132.00 |
| Where the session’s cost goes | ||
| Cache writes | $2.60 | $0.65 |
| Cache reads | $0.40 | $0.20 |
| Uncached input | $0.40 | $0.10 |
| Output | $1.00 | $0.25 |
- caching saves on the session with Claude Opus 5.5 (60%)
- $6.60
- caching saves on the session with Claude Haiku 4.5 (56%)
- $1.55
Why a 4x rate gap becomes 3.7x on a session
Claude Opus 5.5 charges $4 per million input tokens and $20 per million output tokens. Claude Haiku 4.5 charges $1 and $5. Cache writes follow the same 4x: $5 and $8 for Opus 5.5's 5-minute and 1-hour writes, $1.25 and $2 for Haiku 4.5's. Uncached work pays that multiple in full, so the large one-off review costs $0.80 against $0.20 and the output-heavy generation $1.72 against $0.43.
Cache hits break the pattern. Anthropic bills an Opus 5.5 hit at 0.05x its input price, half the 0.1x it charges on Haiku 4.5 and most other Claude models. That makes a hit $0.20 on Opus 5.5 and $0.10 on Haiku 4.5, a 2x gap instead of 4x. The 2M cached tokens in the example session cost $0.40 and $0.20.
Everything else in the session still runs at 4x, so the total lands at 3.7x, a $3.20 difference per session. At 110 sessions a month that is $484.00 against $132.00 at API rates. The largest line on both is cache writes, $2.60 on Opus 5.5 and $0.65 on Haiku 4.5, which is 59% and 54% of each session.
Context, output, and tokenizer differences
Opus 5.5 accepts 1M tokens of context at standard rates, five times Haiku 4.5's 200K, and writes up to 128K tokens of output against 64K. Anthropic names long, sprawling jobs such as codebase-wide migrations and audits as a particular Opus 5.5 strength, and those are the tasks most likely to need more than 200K tokens in a single request.
The token counts themselves aren't directly comparable. Haiku 4.5 uses Anthropic's older tokenizer, and Anthropic notes that Claude models from 4.7 on, Opus 5.5 among them, produce about 30% more tokens for the same text. So the same file costs more tokens on Opus 5.5 as well as more per token, and 200K on Haiku 4.5 holds more text than 200K would on Opus 5.5.
Using Opus 5.5 and Haiku 4.5 together
Anthropic positions the two for different jobs. Opus 5.5 is its recommended starting model for most work, and Anthropic says it performs at the level of Claude Fable 5.1 on most work. Haiku 4.5 is aimed at latency-sensitive work and coding sub-agents, and Anthropic says its coding performance is similar to Claude Sonnet 4 at one-third the cost and more than twice the speed.
That split maps onto multi-agent coding: Opus 5.5 running the main conversation, which it does by default in Claude Code, and Haiku 4.5 handling sub-agents in refactors and migrations, a use Anthropic names for it. Thinking works differently on each. Opus 5.5 has adaptive thinking always on at medium effort by default, while Haiku 4.5 uses manual extended thinking with a token budget.
Haiku 4.5 also has a horizon. Anthropic lists its retirement as not sooner than October 15, 2026, and announced Claude Haiku 5.5 on September 22, 2026, so a sub-agent setup built on Haiku 4.5 now should expect a revisit. To see how a mixed setup splits in practice, EveryToken prices each request in your local Claude Code history at Anthropic's rates and totals cost by model.
Prompt caching
How Anthropic bills cached tokens
Anthropic
Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.
A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.
Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.
In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.
Source: Anthropic: Prompt caching
Your own numbers
See what Claude Opus 5.5 and Claude Haiku 4.5 really cost you.
everyaitoken reads your Claude Code, Cursor, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
How much more does Claude Opus 5.5 cost than Claude Haiku 4.5?
4x per token for input, output, and cache writes, and 2x for cache hits. The example agentic session costs $4.40 on Opus 5.5 and $1.20 on Haiku 4.5, a 3.7x gap.
Can Claude Haiku 4.5 use a 1M token context window?
No. Its context window is 200K tokens, with up to 64K of output. Opus 5.5 takes 1M tokens and writes up to 128K.
Which of the two does Claude Code use by default?
Claude Opus 5.5, on Pro, Max, Team, and Enterprise plans and with an Anthropic API key. Haiku 4.5 is available in Claude Code too, and you can select it.
What is replacing Claude Haiku 4.5?
Claude Haiku 5.5 was announced on September 22, 2026. It lists Haiku 4.5's retirement as not sooner than October 15, 2026.
Sources
- Anthropic: Pricing
- Anthropic docs: Claude Opus 5.5
- Anthropic: Claude Opus 5.5
- Anthropic docs: Fast mode
- Claude Code docs: Model configuration
- Cursor docs: Claude Opus 5.5
- OpenRouter: Claude Opus 5.5
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Anthropic docs: Claude Haiku 4.5
- Anthropic: Introducing Claude Haiku 4.5
- Anthropic: Claude Haiku
- Anthropic docs: Model deprecations
- Cursor docs: Models and pricing
- OpenRouter: Claude Haiku 4.5
- Anthropic: Prompt caching