Skip to content

Model comparison

Claude Sonnet 4.6 vs Gemini 2.5 Pro: prices and limits

Gemini 2.5 Pro costs less than Claude Sonnet 4.6 on every rate, but caps output at 65.5K tokens and now limits who can use it. The full cost and access picture.

· Prices as of September 28, 2026

  • Claude Sonnet 4.6

    Anthropic · Released February 17, 2026 · Previous generation

    The previous Sonnet, pitched at launch as near-Opus capability at the Sonnet price. Anthropic now recommends moving to Claude Sonnet 5.

    Claude Sonnet 4.6 facts and comparisons
  • Gemini 2.5 Pro

    Google · Released June 17, 2025 · Previous generation

    The previous-generation Pro, still stable, and still Gemini CLI's Pro fallback for accounts without preview access.

    Gemini 2.5 Pro facts and comparisons

The short answer

Gemini 2.5 Pro is cheaper on every rate, and the example agentic coding session costs $1.38 on it against $3.60 on Claude Sonnet 4.6. Sonnet 4.6 writes up to 128K tokens of output to Gemini 2.5 Pro's 65.5K and bills its full 1M context at standard rates, while Gemini 2.5 Pro's rates rise above 200K input tokens. Both are previous-generation models, and since September 18, 2026, Google limits Gemini 2.5 Pro to accounts that used it before.

Choose Claude Sonnet 4.6 if

  • You need more than 65.5K tokens of output in one response, since Sonnet 4.6 writes up to 128K.
  • You work in Claude Code or Cursor, and Cursor doesn't list Gemini 2.5 Pro.
  • Your prompts often exceed 200K input tokens and you want one rate across the full 1M window.
  • Your Google account never used Gemini 2.5 Pro, which is now limited to earlier users.

Choose Gemini 2.5 Pro if

  • You want the lower rates: $1.25 input and $10 output per million, against $3 and $15.
  • You use Gemini CLI, which falls back to Gemini 2.5 Pro for accounts without preview access.
  • Your account already used Gemini 2.5 Pro, so Google's access limit doesn't affect you.

Side by side

Specs and prices

FactClaude Sonnet 4.6Gemini 2.5 Pro
MakerAnthropicGoogle
API model idclaude-sonnet-4-6gemini-2.5-pro
ReleasedFebruary 17, 2026June 17, 2025
StatusPrevious generationPrevious generation
Context window1M tokens1.05M tokens
Max output128K tokens65.5K tokens
Open weightsNoNo
Input, per 1M tokens$3$1.25
Cache hit, per 1M$0.30$0.125
Cache write, per 1M$3.75 (5-minute), $6 (1-hour)$1.25 (same as input)
Output, per 1M tokens$15$10
Runs inClaude Code, Cursor, OpenCode, and OpenRouterGemini CLI, OpenCode, and OpenRouter

Standard API rates in US dollars, as published by each maker (Claude Sonnet 4.6: September 26, 2026; Gemini 2.5 Pro: September 28, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Claude Sonnet 4.6: The full 1M context window is billed at standard rates on the API. Gemini 2.5 Pro: Prompts over 200K input tokens cost $2.50 input, $0.25 cached, and $15 output per million tokens.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadClaude Sonnet 4.6Gemini 2.5 Pro
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$3.60$1.38
Large one-off review, 150K input with no cache hits, 10K output$0.60$0.29
Output-heavy generation, 30K input, 80K output$1.29$0.84
A month of sessions, 110 sessions: 5 a day, 22 working days$396.00$151.25
Where the session’s cost goes
Cache writes$1.95$0.50
Cache reads$0.60$0.25
Uncached input$0.30$0.13
Output$0.75$0.50
caching saves on the session with Claude Sonnet 4.6 (56%)
$4.65
caching saves on the session with Gemini 2.5 Pro (62%)
$2.25

How much less does Gemini 2.5 Pro cost?

Gemini 2.5 Pro charges $1.25 per million input tokens, $0.125 per million cached, and $10 per million output. Claude Sonnet 4.6 asks $3 input, $0.30 cached, and $15 output. Both bill a cache hit at 10% of input, so the input and hit gaps are the same 2.4x, and output is 1.5x apart.

Cache writes widen the gap. Google publishes no separate cache-write price, so written tokens are priced as ordinary input at $1.25, while Anthropic bills $3.75 for a 5-minute write and $6 for a 1-hour write. In the example session, writes cost $1.95 on Sonnet 4.6 and $0.50 on Gemini 2.5 Pro, and the session comes to $3.60 against $1.38, 2.6x. Uncached work shows smaller gaps: 2.1x on the large review, $0.60 against $0.29, and 1.5x on the output-heavy generation, $1.29 against $0.84.

At 110 sessions a month, the estimate is $396.00 on Sonnet 4.6 and $151.25 on Gemini 2.5 Pro at API rates. Google's explicit caching adds a storage charge of $4.50 per million tokens per hour on Pro models, which the tables leave out. Implicit caching, on by default, applies the discount automatically when a request repeats a prefix Google has cached, but a hit isn't assured.

Output limits and the 200K price tier

The biggest spec difference is output. Sonnet 4.6 can write up to 128K tokens in one response, and Gemini 2.5 Pro up to 65.5K. Both accept about a million tokens of context: 1M on Sonnet 4.6 and 1.05M on Gemini 2.5 Pro.

Long prompts are priced differently. Gemini 2.5 Pro charges $2.50 input, $0.25 cached, and $15 output per million once a prompt goes over 200K input tokens. Sonnet 4.6 bills its full 1M window at standard rates on the API. Past 200K, Gemini 2.5 Pro's output rate matches Sonnet 4.6's $15, while its input stays below $3. The example session keeps every request under 200K, so the tables use the standard tier.

Thinking is configured differently too. Gemini 2.5 Pro uses thinking budgets rather than thinking levels, and Gemini CLI sends a budget of 8,192 tokens. The two makers also tokenize text differently, so a prompt's token count won't match across them, while the fixed workloads here assume it does.

Access, tools, and what replaced each model

Since September 18, 2026, Google limits Gemini 2.5 Pro to accounts that used it before. It is not deprecated, and it is still Gemini CLI's Pro fallback for accounts without preview access. It runs in Gemini CLI, OpenRouter, and OpenCode, but Cursor and GitHub Copilot don't list it. Google's current Pro model is Gemini 3.1 Pro Preview, still in preview.

Sonnet 4.6 is a legacy model that stays available on the Anthropic API, and runs in Claude Code, Cursor, OpenRouter, and OpenCode. In GitHub Copilot it was retired on September 1, 2026, with an exception for annual Copilot Pro and Pro+ subscribers. Anthropic's suggested replacement is Claude Sonnet 5.

At launch, Anthropic pitched Sonnet 4.6 as approaching Opus-level intelligence at a more practical price. Google describes Gemini 2.5 Pro as "A Pro model which excels at coding and complex reasoning tasks," and names analyzing large codebases and documents with long context among its strengths.

Prompt caching

How each maker bills cached tokens

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

Google

Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.

Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.

On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.

Source: Google: Context caching

Your own numbers

See what Claude Sonnet 4.6 and Gemini 2.5 Pro really cost you.

everyaitoken reads your Claude Code, Cursor, OpenCode, OpenRouter, and Gemini CLI history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is Gemini 2.5 Pro cheaper than Claude Sonnet 4.6?

Yes, on every rate. The example agentic session costs $1.38 on Gemini 2.5 Pro and $3.60 on Sonnet 4.6. The gap is smallest on output-heavy work, where it is 1.5x.

Can anyone still use Gemini 2.5 Pro?

Not since September 18, 2026. Google now limits it to accounts that used it before, though the model is not deprecated.

Does Gemini 2.5 Pro have a lower output limit than Claude Sonnet 4.6?

Yes. Gemini 2.5 Pro writes up to 65.5K tokens per response, and Claude Sonnet 4.6 up to 128K. For long generated files, that limit may decide the choice before price does.

How do I compare my Claude Code and Gemini CLI spend?

EveryToken reads local Claude Code and Gemini CLI history on your Mac and prices every request at API rates, so the two show side by side with what caching saved on each.

  • Claude Sonnet 5 vs Claude Sonnet 4.6

    Claude Sonnet 5 costs 33% less per token than Claude Sonnet 4.6, but its tokenizer counts up to 1.35x as many tokens. What the upgrade saves in practice.

  • Gemini 2.5 Pro vs Gemini 2.5 Flash

    Gemini 2.5 Pro costs about 4x Gemini 2.5 Flash, and since September 18, 2026, only earlier users can reach either. What each costs and where to go next.

  • Gemini 3.1 Pro Preview vs Gemini 2.5 Pro

    Gemini 3.1 Pro Preview costs 60% more per input token than Gemini 2.5 Pro and is still a preview, while 2.5 Pro now limits new access. How to choose.

  • Claude Sonnet 4.6 vs GPT-5.3-Codex

    Claude Sonnet 4.6 and GPT-5.3-Codex launched 12 days apart in February 2026. One takes 1M tokens of context, the other costs 46% less per session. Tradeoffs.

  • Claude Sonnet 4.6 vs GPT-5.4

    Claude Sonnet 4.6 and GPT-5.4 both charge $15 per million output tokens, so the cost gap lives in cache writes. What that means for sessions and upgrades.

  • Claude Fable 5.1 vs Claude Fable 5

    Claude Fable 5.1 keeps Claude Fable 5's rates but cuts a cache hit from $1 to $0.25 per million. What that saves on agentic sessions, and what it doesn't.