Model comparison
Claude Sonnet 5 or Claude Haiku 4.5 for coding sub-agents?
Claude Haiku 4.5 costs half as much as Claude Sonnet 5, with a 200K context window and a retirement date ahead. When the cheaper model fits a coding workflow.
· Prices as of September 26, 2026
Claude Sonnet 5
Anthropic · Released June 30, 2026
Anthropic's balance of speed and intelligence, and a drop-in upgrade for Claude Sonnet 4.6. Anthropic pitches it as close to Claude Opus 4.8 at a lower price.
Claude Sonnet 5 facts and comparisonsClaude Haiku 4.5
Anthropic · Released October 15, 2025
Anthropic's fastest and cheapest current model, aimed at latency-sensitive work and at coding sub-agents.
Claude Haiku 4.5 facts and comparisons
The short answer
Claude Haiku 4.5 costs half as much as Claude Sonnet 5 on every rate, so the example agentic coding session is $1.20 against $2.40. Sonnet 5 has a 1M context window to Haiku 4.5's 200K and writes up to 128K tokens of output to its 64K. Haiku 4.5 suits latency-sensitive work and sub-agents with small contexts, with the caveat that Anthropic lists its retirement as not sooner than October 15, 2026.
Choose Claude Sonnet 5 if
- Your tasks need more than 200K tokens of context, such as reading across a large codebase in one pass.
- You want the model Anthropic pitches as close to Claude Opus 4.8 at a lower price, with agentic planning and tool use.
- You want adaptive thinking and effort levels rather than a fixed thinking budget.
- You'd rather not plan a migration soon: Anthropic lists Haiku 4.5's retirement as not sooner than October 15, 2026.
Choose Claude Haiku 4.5 if
- You run sub-agents in multi-agent refactors and migrations, a use Anthropic names for Haiku 4.5.
- Your work is latency-sensitive and each task fits within 200K tokens of context and 64K of output.
- You want the lower per-token rates in this pair: $1 input and $5 output per million.
Side by side
Specs and prices
| Fact | Claude Sonnet 5 | Claude Haiku 4.5 |
|---|---|---|
| Maker | Anthropic | Anthropic |
| API model id | claude-sonnet-5 | claude-haiku-4-5 |
| Released | June 30, 2026 | October 15, 2025 |
| Status | Current | Current |
| Context window | 1M tokens | 200K tokens |
| Max output | 128K tokens | 64K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $2 | $1 |
| Cache hit, per 1M | $0.20 | $0.10 |
| Cache write, per 1M | $2.50 (5-minute), $4 (1-hour) | $1.25 (5-minute), $2 (1-hour) |
| Output, per 1M tokens | $10 | $5 |
| Runs in | Claude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot | Claude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker on September 26, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Claude Sonnet 5: The full 1M context window is billed at standard rates. The launch price became the standard price on August 10, 2026, and a planned increase was cancelled.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Claude Sonnet 5 | Claude Haiku 4.5 |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $2.40 | $1.20 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.40 | $0.20 |
| Output-heavy generation, 30K input, 80K output | $0.86 | $0.43 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $264.00 | $132.00 |
| Where the session’s cost goes | ||
| Cache writes | $1.30 | $0.65 |
| Cache reads | $0.40 | $0.20 |
| Uncached input | $0.20 | $0.10 |
| Output | $0.50 | $0.25 |
- caching saves on the session with Claude Sonnet 5 (56%)
- $3.10
- caching saves on the session with Claude Haiku 4.5 (56%)
- $1.55
Is Claude Haiku 4.5 half the price of Claude Sonnet 5?
On every line of the rate card. Claude Haiku 4.5 charges $1 per million input tokens, $5 per million output tokens, $1.25 and $2 for 5-minute and 1-hour cache writes, and $0.10 for a cache hit. Claude Sonnet 5 charges exactly twice each of those: $2 input, $10 output, $2.50 and $4 for cache writes, and $0.20 for a hit. Both bill hits at 0.1x input, so caching doesn't shift the ratio.
Every workload therefore lands at 2x. The example session costs $2.40 against $1.20, the large one-off review $0.40 against $0.20, and the output-heavy generation $0.86 against $0.43. A month of 110 sessions comes to $264.00 on Sonnet 5 and $132.00 on Haiku 4.5 at API rates, with caching saving 56% on each.
For the same text, the gap is wider than the rate card shows. Haiku 4.5 uses Anthropic's older tokenizer, and Anthropic notes that Claude models from 4.7 on produce about 30% more tokens for the same text. A file that counts as a given number of tokens on Haiku 4.5 will usually count as more on Sonnet 5, so the real cost gap for identical work is usually above 2x.
Where Haiku 4.5's 200K context window matters
Haiku 4.5 accepts 200K tokens of context, a fifth of Sonnet 5's 1M, and writes up to 64K tokens of output against 128K. Each request in the example session stays below 200K input tokens, which both models can take. An agent that loads a large codebase, a long log, or a full conversation history into one request can reach Haiku 4.5's ceiling where Sonnet 5 still has room.
That is why the two can work together rather than compete. Anthropic names powering sub-agents in multi-agent refactors and migrations as a Haiku 4.5 use. In that kind of setup, one model holds the whole task while sub-agents take narrow pieces with small contexts, which is where Haiku 4.5's window is least likely to bind and its lower rates add up.
Thinking, effort, and Haiku 4.5's retirement date
The two models handle reasoning differently. Sonnet 5 thinks adaptively by default and starts at high effort. Haiku 4.5 uses manual extended thinking with a token budget you set, and has no adaptive thinking or effort levels. A fixed budget makes Haiku 4.5's thinking spend easier to cap, while adaptive thinking leaves more of that decision to Sonnet 5.
Anthropic positions Haiku 4.5 for latency-sensitive work and coding sub-agents, and says it delivers coding performance similar to Claude Sonnet 4 at one-third the cost and more than twice the speed. That comparison is with Sonnet 4, not Sonnet 5. For Sonnet 5, Anthropic claims performance close to Claude Opus 4.8 at lower prices.
Plan for change on the Haiku side. By Anthropic's listing, Haiku 4.5 retires no sooner than October 15, 2026, and Claude Haiku 5.5 was announced on September 22, 2026. Both current models are listed in Claude Code, Cursor, OpenRouter, OpenCode, and GitHub Copilot.
Prompt caching
How Anthropic bills cached tokens
Anthropic
Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.
A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.
Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.
In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.
Source: Anthropic: Prompt caching
Your own numbers
See what Claude Sonnet 5 and Claude Haiku 4.5 really cost you.
everyaitoken reads your Claude Code, Cursor, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
What does Anthropic say Claude Haiku 4.5 is for?
Latency-sensitive work and coding sub-agents, including sub-agents in multi-agent refactors and migrations. Anthropic says its coding performance is similar to Claude Sonnet 4 at one-third the cost.
Can Claude Haiku 4.5 handle a 1M token context?
No. Haiku 4.5's context window is 200K tokens and its maximum output is 64K. Sonnet 5 accepts 1M tokens and writes up to 128K.
Is Claude Haiku 4.5 twice as cheap as Claude Sonnet 5 for the same work?
At least. Each rate is half, and Haiku 4.5's older tokenizer counts fewer tokens for the same text than Sonnet 5's, so identical work usually costs less than half as much on Haiku 4.5.
How do I see what my sub-agents cost next to the main model?
EveryToken reads your local Claude Code history and prices every request at API rates, split by model. That shows how much of your spend goes to Haiku 4.5 and how much to the model running the main conversation.
Sources
- Anthropic: Pricing
- Anthropic docs: Claude Sonnet 5
- Anthropic: Introducing Claude Sonnet 5
- Claude Code docs: Model configuration
- Cursor docs: Claude Sonnet 5
- OpenRouter: Claude Sonnet 5
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Anthropic docs: Claude Haiku 4.5
- Anthropic: Introducing Claude Haiku 4.5
- Anthropic: Claude Haiku
- Anthropic docs: Model deprecations
- Cursor docs: Models and pricing
- OpenRouter: Claude Haiku 4.5
- Anthropic: Prompt caching