Skip to content

Anthropic

Claude Sonnet 5: price, context window, and caching

Anthropic's balance of speed and intelligence, and a drop-in upgrade for Claude Sonnet 4.6. Anthropic pitches it as close to Claude Opus 4.8 at a lower price.

Released June 30, 2026 · Prices as of September 26, 2026

In Anthropic’s words

“Claude Sonnet 5 is built to be the most agentic Sonnet model yet.”

Anthropic: Introducing Claude Sonnet 5

What Anthropic says it’s good at

  • Agentic work: planning, and using tools like browsers and terminals on its own Source
  • Performance close to Claude Opus 4.8 at lower prices Source
  • Better reasoning, tool use, and coding than Claude Sonnet 4.6 Source

Facts

Specs and prices

FactClaude Sonnet 5
MakerAnthropic
API model idclaude-sonnet-5
ReleasedJune 30, 2026
StatusCurrent
Context window1M tokens
Max output128K tokens
Open weightsNo
Input, per 1M tokens$2
Cache hit, per 1M$0.20
Cache write, per 1M$2.50 (5-minute), $4 (1-hour)
Output, per 1M tokens$10
Runs inClaude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by the maker on September 26, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Claude Sonnet 5: The full 1M context window is billed at standard rates. The launch price became the standard price on August 10, 2026, and a planned increase was cancelled.

Good to know

  • Adaptive thinking is on by default, and the default effort is high.
  • The sonnet alias in Claude Code resolves to it on the Anthropic API.
  • Its newer tokenizer counts about 1.0 to 1.35x as many tokens as Claude Sonnet 4.6 for the same text.

Cost

What typical work costs

Example token counts at Claude Sonnet 5’s published rates. On the agentic session, caching saves $3.10 against billing every token as ordinary input.

Example workload costs for Claude Sonnet 5
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$2.40
Large one-off review, 150K input with no cache hits, 10K output$0.40
Output-heavy generation, 30K input, 80K output$0.86
A month of sessions, 110 sessions: 5 a day, 22 working days$264.00

Prompt caching

How Anthropic bills cached tokens

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

Claude Sonnet 5 compared

  • Claude Fable 5.1 vs Claude Sonnet 5

    Claude Fable 5.1 lists at 5x the rates of Claude Sonnet 5, but cache hits cost $0.25 against $0.20. Why sessions cost 4.4x, and who Anthropic aims Fable 5.1 at.

  • Claude Opus 5.5 vs Claude Sonnet 5

    Claude Opus 5.5 costs twice as much per token as Claude Sonnet 5, yet cache hits cost $0.20 on both. What that means for agentic coding sessions.

  • Claude Sonnet 5 vs Claude Haiku 4.5

    Claude Haiku 4.5 costs half as much as Claude Sonnet 5, with a 200K context window and a retirement date ahead. When the cheaper model fits a coding workflow.

  • GLM-5.3-Flash vs Claude Sonnet 5

    GLM-5.3-Flash costs $0.16 for a cached coding session that costs $2.40 on Claude Sonnet 5. What drives a 15x gap, and what price alone can't tell you.

  • Claude Sonnet 5 vs Claude Sonnet 4.6

    Claude Sonnet 5 costs 33% less per token than Claude Sonnet 4.6, but its tokenizer counts up to 1.35x as many tokens. What the upgrade saves in practice.

  • Claude Sonnet 5 vs Gemini 3.1 Pro Preview

    Claude Sonnet 5 and Gemini 3.1 Pro Preview both charge $2 input, but Gemini's rates rise past 200K tokens and it is still a preview. Session: $2.40 vs $2.00.

  • Claude Sonnet 5 vs Gemini 3.5 Flash

    Gemini 3.5 Flash, now Google's legacy Flash, costs $1.50 against $2.40 on Claude Sonnet 5 per cached coding session. Most of the gap is cache writes.

  • Claude Sonnet 5 vs Gemini 3.8 Flash

    Gemini 3.8 Flash runs a cached coding session for $0.71 against $2.40 on Claude Sonnet 5, on introductory rates that end December 31, 2026.

  • Claude Sonnet 5 vs GPT-5.5

    GPT-5.5 costs 2.5x Claude Sonnet 5 for input and 3x for output, and it leaves ChatGPT and Codex sign-in on October 14, 2026. What that means for coding.

  • Claude Sonnet 5 vs GPT-5.6 Sol

    GPT-5.6 Sol charges twice Claude Sonnet 5's rates on promotional pricing that runs through at least November 21, 2026. A cached session: $4.20 vs $2.40.

  • Claude Sonnet 5 vs GPT-5.6 Terra

    Claude Sonnet 5 and GPT-5.6 Terra share a $2 input rate. Terra costs less on cached sessions, Sonnet 5 on output-heavy work. The numbers, line by line.

  • Claude Sonnet 5 vs GPT-6 Astra

    GPT-6 Astra charges 5x Claude Sonnet 5's rates on every kind of token. One cached coding session costs $10.50 against $2.40, a 4.4x gap. Here is why.

  • Claude Sonnet 5 vs GPT-6 Sol

    Claude Sonnet 5 and GPT-6 Sol list the same $2 input and $10 output rates. Cache writes decide the gap: $2.40 against $2.10 for one coding session.

  • DeepSeek-V4.1-Flash vs Claude Sonnet 5

    A cached coding session costs $0.22 on DeepSeek-V4.1-Flash and $2.40 on Claude Sonnet 5. Where the 10.9x gap comes from, and what open weights change.

  • GLM-5.3 vs Claude Sonnet 5

    GLM-5.3 undercuts Claude Sonnet 5 by 40% on an agentic coding session, $1.44 against $2.40. Most of it comes from Anthropic's cache-write premium.

  • Grok 4.7 vs Claude Sonnet 5

    Grok 4.7 and Claude Sonnet 5 both charge $2 per million input tokens. A cached coding session costs $2.30 against $2.40, but output-heavy work splits wider.

  • Grok Build 0.1 vs Claude Sonnet 5

    Grok Build 0.1 costs $1.00 per cached coding session against $2.40 on Claude Sonnet 5, but it is in early access, with a 256K window and a price step at 200K.

  • DeepSeek-V4-Pro vs Claude Sonnet 5

    DeepSeek-V4-Pro costs $0.95 for a cached coding session that costs $2.40 on Claude Sonnet 5, and cache writes explain most of it. Rates, limits, and tools.

  • Kimi K3 vs Claude Sonnet 5

    Kimi K3 lists 1.5x the rates of Claude Sonnet 5, yet an agentic coding session costs just $0.45 more on it, $2.85 against $2.40. Cache writes explain why.

  • MiniMax M3 vs Claude Sonnet 5

    MiniMax M3 costs $0.33 per cached coding session against $2.40 on Claude Sonnet 5. Both have a 1M window, and they differ on tools, caching, and long prompts.

  • Mistral Medium 3.5 vs Claude Sonnet 5

    Mistral Medium 3.5 lists rates 25% below Claude Sonnet 5, yet a cached coding session costs 40% less, $1.43 against $2.40. Cache writes explain the difference.

  • Qwen3.8-Max vs Claude Sonnet 5

    Qwen3.8-Max and Claude Sonnet 5 both charge $2 per million input tokens. Output and cache writes set them apart: $1.80 against $2.40 per coding session.

Your own numbers

See what Claude Sonnet 5 really costs you.

everyaitoken reads your Claude Code, Cursor, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math