Skip to content

Model comparison

Claude Sonnet 5 vs Claude Sonnet 4.6: price and tokenizer

Claude Sonnet 5 costs 33% less per token than Claude Sonnet 4.6, but its tokenizer counts up to 1.35x as many tokens. What the upgrade saves in practice.

· Prices as of September 26, 2026

  • Claude Sonnet 5

    Anthropic · Released June 30, 2026

    Anthropic's balance of speed and intelligence, and a drop-in upgrade for Claude Sonnet 4.6. Anthropic pitches it as close to Claude Opus 4.8 at a lower price.

    Claude Sonnet 5 facts and comparisons
  • Claude Sonnet 4.6

    Anthropic · Released February 17, 2026 · Previous generation

    The previous Sonnet, pitched at launch as near-Opus capability at the Sonnet price. Anthropic now recommends moving to Claude Sonnet 5.

    Claude Sonnet 4.6 facts and comparisons

The short answer

Claude Sonnet 5 lists at $2 input and $10 output per million tokens, 33% below Claude Sonnet 4.6, and Anthropic calls it a drop-in upgrade. The example agentic coding session costs $2.40 on Sonnet 5 against $3.60, though Sonnet 5's newer tokenizer counts about 1.0 to 1.35x as many tokens for the same text, which eats into that gap. Sonnet 4.6 is now a legacy model, and Anthropic recommends the move.

Choose Claude Sonnet 5 if

  • You want the lower rates: $2 input, $10 output, and $0.20 per cache hit per million, against $3 input, $15 output, and $0.30 per hit.
  • You use the sonnet alias in Claude Code, which resolves to Sonnet 5 on the Anthropic API.
  • You want the gains Anthropic claims over Sonnet 4.6 in reasoning, tool use, and coding, and in agentic work like planning and running browsers and terminals.
  • You use GitHub Copilot, which retired Sonnet 4.6 on September 1, 2026, except for annual Copilot Pro and Pro+ subscribers.

Choose Claude Sonnet 4.6 if

  • You have prompts, evaluations, or token budgets calibrated to Sonnet 4.6's older tokenizer and aren't ready to re-measure them.
  • You need the model to stay fixed while you test the new one, and Sonnet 4.6 remains available on the API as a legacy model.

Side by side

Specs and prices

FactClaude Sonnet 5Claude Sonnet 4.6
MakerAnthropicAnthropic
API model idclaude-sonnet-5claude-sonnet-4-6
ReleasedJune 30, 2026February 17, 2026
StatusCurrentPrevious generation
Context window1M tokens1M tokens
Max output128K tokens128K tokens
Open weightsNoNo
Input, per 1M tokens$2$3
Cache hit, per 1M$0.20$0.30
Cache write, per 1M$2.50 (5-minute), $4 (1-hour)$3.75 (5-minute), $6 (1-hour)
Output, per 1M tokens$10$15
Runs inClaude Code, Cursor, OpenCode, OpenRouter, and GitHub CopilotClaude Code, Cursor, OpenCode, and OpenRouter

Standard API rates in US dollars, as published by each maker on September 26, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Claude Sonnet 5: The full 1M context window is billed at standard rates. The launch price became the standard price on August 10, 2026, and a planned increase was cancelled. Claude Sonnet 4.6: The full 1M context window is billed at standard rates on the API.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadClaude Sonnet 5Claude Sonnet 4.6
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$2.40$3.60
Large one-off review, 150K input with no cache hits, 10K output$0.40$0.60
Output-heavy generation, 30K input, 80K output$0.86$1.29
A month of sessions, 110 sessions: 5 a day, 22 working days$264.00$396.00
Where the session’s cost goes
Cache writes$1.30$1.95
Cache reads$0.40$0.60
Uncached input$0.20$0.30
Output$0.50$0.75
caching saves on the session with Claude Sonnet 5 (56%)
$3.10
caching saves on the session with Claude Sonnet 4.6 (56%)
$4.65

How much cheaper is Claude Sonnet 5 per token?

Every rate dropped by a third. Claude Sonnet 5 charges $2 per million input tokens and $10 per million output tokens, against $3 and $15 for Claude Sonnet 4.6. Cache writes fall from $3.75 and $6 to $2.50 and $4 for 5-minute and 1-hour writes, and cache hits from $0.30 to $0.20. Both models bill hits at 0.1x input, so the discount structure is unchanged.

Because every line moved by the same ratio, every workload moves by it too. The example session costs $2.40 against $3.60, the large one-off review $0.40 against $0.60, and the output-heavy generation $0.86 against $1.29, 33% less each time. At 110 sessions a month that is $264.00 against $396.00, a $132.00 difference at API rates.

Sonnet 5's price has a short history. It launched at these rates, which became the standard price on August 10, 2026, when Anthropic cancelled a planned increase.

Does Claude Sonnet 5 use more tokens than Sonnet 4.6?

For the same text, it can. Anthropic says Sonnet 5's newer tokenizer counts about 1.0 to 1.35x as many tokens as Sonnet 4.6 for the same text. Sonnet 4.6 uses Anthropic's older tokenizer, and Anthropic's broader note is that Claude models from 4.7 on produce about 30% more tokens for the same text.

That changes the real saving. At the low end of the range, text costs 33% less on Sonnet 5, as the tables show. At the top end, Sonnet 5 still comes out cheaper for the same text, because 1.35x as many tokens is still less than the 1.5x price ratio, but by a much smaller margin. Where your own prompts land in that range depends on what they contain, so code, prose, and logs may each shift by a different amount.

Token budgets are the other thing to recheck. Limits you set in tokens, such as a maximum output length, hold less text on Sonnet 5. Both models accept 1M tokens of context and write up to 128K of output, and Sonnet 4.6 bills its full 1M at standard rates on the API.

What Anthropic says changed in Claude Sonnet 5

Anthropic positions Sonnet 5 as its balance of speed and intelligence and a drop-in upgrade for Sonnet 4.6. It claims better reasoning, tool use, and coding than Sonnet 4.6, performance close to Claude Opus 4.8 at lower prices, and stronger agentic work: planning, and using tools like browsers and terminals on its own. In Anthropic's words, "Claude Sonnet 5 is built to be the most agentic Sonnet model yet."

Sonnet 5 has adaptive thinking on by default and starts at high effort. Thinking adds tokens, so a switch can change how much each response writes as well as the rate it is billed at. Match settings before reading much into a cost change.

Sonnet 4.6 is a legacy model that stays available on the API, and it remains in Claude Code, Cursor, OpenRouter, and OpenCode. Copilot users lost it on September 1, 2026, apart from annual Copilot Pro and Pro+ subscribers. In Claude Code, the sonnet alias now resolves to Sonnet 5 on the Anthropic API.

Prompt caching

How Anthropic bills cached tokens

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

Your own numbers

See what Claude Sonnet 5 and Claude Sonnet 4.6 really cost you.

everyaitoken reads your Claude Code, Cursor, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is Claude Sonnet 5 cheaper than Claude Sonnet 4.6?

Per token, yes: 33% less on every rate. Its tokenizer can count up to 1.35x as many tokens for the same text, so the saving on a given prompt ranges from the full 33% down to a much smaller margin.

Is Claude Sonnet 5 a drop-in replacement for Claude Sonnet 4.6?

Anthropic calls it one. Both have a 1M context window and 128K of output. Recheck any limits set in tokens, since the newer tokenizer counts the same text differently.

Is Claude Sonnet 4.6 being retired?

Anthropic keeps it as a legacy model on the API and recommends Sonnet 5. GitHub Copilot retired it on September 1, 2026, except for annual Copilot Pro and Pro+ subscribers.

How do I measure the change on my own work?

EveryToken reads your Claude Code history on your Mac, prices each request at Anthropic's rates, and splits cost by model. A before-and-after comparison shows the net effect of the lower price and the newer tokenizer together.

  • Claude Fable 5.1 vs Claude Sonnet 5

    Claude Fable 5.1 lists at 5x the rates of Claude Sonnet 5, but cache hits cost $0.25 against $0.20. Why sessions cost 4.4x, and who Anthropic aims Fable 5.1 at.

  • Claude Opus 5.5 vs Claude Sonnet 5

    Claude Opus 5.5 costs twice as much per token as Claude Sonnet 5, yet cache hits cost $0.20 on both. What that means for agentic coding sessions.

  • Claude Sonnet 5 vs Claude Haiku 4.5

    Claude Haiku 4.5 costs half as much as Claude Sonnet 5, with a 200K context window and a retirement date ahead. When the cheaper model fits a coding workflow.

  • Claude Sonnet 4.6 vs GPT-5.3-Codex

    Claude Sonnet 4.6 and GPT-5.3-Codex launched 12 days apart in February 2026. One takes 1M tokens of context, the other costs 46% less per session. Tradeoffs.

  • Claude Sonnet 4.6 vs Gemini 2.5 Pro

    Gemini 2.5 Pro costs less than Claude Sonnet 4.6 on every rate, but caps output at 65.5K tokens and now limits who can use it. The full cost and access picture.

  • Claude Sonnet 4.6 vs GPT-5.4

    Claude Sonnet 4.6 and GPT-5.4 both charge $15 per million output tokens, so the cost gap lives in cache writes. What that means for sessions and upgrades.