Skip to content

Anthropic

Claude Sonnet 4.6: price, context window, and caching

The previous Sonnet, pitched at launch as near-Opus capability at the Sonnet price. Anthropic now recommends moving to Claude Sonnet 5.

Released February 17, 2026 · Prices as of September 26, 2026

In Anthropic’s words

“It's a full upgrade of the model's skills across coding, computer use, long-context reasoning, agent planning, knowledge work, and design.”

Anthropic: Introducing Claude Sonnet 4.6

What Anthropic says it’s good at

  • Approaching Opus-level intelligence at a more practical price Source
  • Preferred over Claude Sonnet 4.5 by early Claude Code testers Source

Facts

Specs and prices

FactClaude Sonnet 4.6
MakerAnthropic
API model idclaude-sonnet-4-6
ReleasedFebruary 17, 2026
StatusPrevious generation
Context window1M tokens
Max output128K tokens
Open weightsNo
Input, per 1M tokens$3
Cache hit, per 1M$0.30
Cache write, per 1M$3.75 (5-minute), $6 (1-hour)
Output, per 1M tokens$15
Runs inClaude Code, Cursor, OpenCode, and OpenRouter

Standard API rates in US dollars, as published by the maker on September 26, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Claude Sonnet 4.6: The full 1M context window is billed at standard rates on the API.

Good to know

  • A legacy model that stays available on the API. GitHub Copilot retired it on September 1, 2026, except for annual Copilot Pro and Pro+ subscribers.
  • Anthropic lists its retirement as not sooner than February 17, 2027.
  • Effort runs low, medium, high, or max. The default is high, and Anthropic recommends medium for agentic coding.
  • It uses Anthropic's older tokenizer, so the same text counts as fewer tokens than on Claude models from 4.7 on, which produce about 30% more.

Cost

What typical work costs

Example token counts at Claude Sonnet 4.6’s published rates. On the agentic session, caching saves $4.65 against billing every token as ordinary input.

Example workload costs for Claude Sonnet 4.6
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$3.60
Large one-off review, 150K input with no cache hits, 10K output$0.60
Output-heavy generation, 30K input, 80K output$1.29
A month of sessions, 110 sessions: 5 a day, 22 working days$396.00

Prompt caching

How Anthropic bills cached tokens

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

Claude Sonnet 4.6 compared

  • Claude Sonnet 4.6 vs GPT-5.3-Codex

    Claude Sonnet 4.6 and GPT-5.3-Codex launched 12 days apart in February 2026. One takes 1M tokens of context, the other costs 46% less per session. Tradeoffs.

  • Claude Sonnet 4.6 vs Gemini 2.5 Pro

    Gemini 2.5 Pro costs less than Claude Sonnet 4.6 on every rate, but caps output at 65.5K tokens and now limits who can use it. The full cost and access picture.

  • Claude Sonnet 4.6 vs GPT-5.4

    Claude Sonnet 4.6 and GPT-5.4 both charge $15 per million output tokens, so the cost gap lives in cache writes. What that means for sessions and upgrades.

  • Claude Sonnet 5 vs Claude Sonnet 4.6

    Claude Sonnet 5 costs 33% less per token than Claude Sonnet 4.6, but its tokenizer counts up to 1.35x as many tokens. What the upgrade saves in practice.

Your own numbers

See what Claude Sonnet 4.6 really costs you.

everyaitoken reads your Claude Code, Cursor, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math