Skip to content

Model comparison

Claude Sonnet 5 vs Gemini 3.1 Pro Preview for coding costs

Claude Sonnet 5 and Gemini 3.1 Pro Preview both charge $2 input, but Gemini's rates rise past 200K tokens and it is still a preview. Session: $2.40 vs $2.00.

· Prices as of September 28, 2026

  • Claude Sonnet 5

    Anthropic · Released June 30, 2026

    Anthropic's balance of speed and intelligence, and a drop-in upgrade for Claude Sonnet 4.6. Anthropic pitches it as close to Claude Opus 4.8 at a lower price.

    Claude Sonnet 5 facts and comparisons
  • Gemini 3.1 Pro Preview

    Google · Released February 19, 2026 · Preview

    Google's only current Pro model, still in preview, positioned for deep reasoning and agentic coding.

    Gemini 3.1 Pro Preview facts and comparisons

The short answer

Gemini 3.1 Pro Preview costs less on the example agentic coding session, $2.00 against $2.40 for Claude Sonnet 5, because Google bills cache writes as ordinary input. Sonnet 5 costs less on uncached and output-heavy work and can write 128K tokens in one response against 65.5K. Gemini 3.1 Pro Preview is still a preview, its rates rise on prompts over 200K input tokens, and GitHub Copilot has already dropped it.

Choose Claude Sonnet 5 if

  • Your work is output-heavy: Sonnet 5 charges $10 per million output tokens against $12 per million on Gemini, and the output-heavy generation costs $0.86 against $1.02.
  • You need responses longer than 65.5K tokens, since Sonnet 5 writes up to 128K.
  • You use GitHub Copilot, which retired Gemini 3.1 Pro Preview on September 1, 2026, and still offers Claude Sonnet 5.
  • You want a current, non-preview model for production agents in Claude Code.

Choose Gemini 3.1 Pro Preview if

  • Your sessions write a lot of context to the cache, where Google's lack of a write premium makes the example session $0.40 cheaper.
  • You work in Gemini CLI, whose default auto model uses Gemini 3.1 Pro Preview as its Pro half.
  • Your agent mixes custom tools with bash, and you want the customtools endpoint Google says is better at prioritizing them, at the same price.
  • Your prompts stay under 200K input tokens, below the point where its rates rise.

Side by side

Specs and prices

FactClaude Sonnet 5Gemini 3.1 Pro Preview
MakerAnthropicGoogle
API model idclaude-sonnet-5gemini-3.1-pro-preview
ReleasedJune 30, 2026February 19, 2026
StatusCurrentPreview
Context window1M tokens1.05M tokens
Max output128K tokens65.5K tokens
Open weightsNoNo
Input, per 1M tokens$2$2
Cache hit, per 1M$0.20$0.20
Cache write, per 1M$2.50 (5-minute), $4 (1-hour)$2 (same as input)
Output, per 1M tokens$10$12
Runs inClaude Code, Cursor, OpenCode, OpenRouter, and GitHub CopilotCursor, Gemini CLI, OpenCode, and OpenRouter

Standard API rates in US dollars, as published by each maker (Claude Sonnet 5: September 26, 2026; Gemini 3.1 Pro Preview: September 28, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Claude Sonnet 5: The full 1M context window is billed at standard rates. The launch price became the standard price on August 10, 2026, and a planned increase was cancelled. Gemini 3.1 Pro Preview: Prompts over 200K input tokens cost $4 input, $0.40 cached, and $18 output per million tokens.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadClaude Sonnet 5Gemini 3.1 Pro Preview
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$2.40$2.00
Large one-off review, 150K input with no cache hits, 10K output$0.40$0.42
Output-heavy generation, 30K input, 80K output$0.86$1.02
A month of sessions, 110 sessions: 5 a day, 22 working days$264.00$220.00
Where the session’s cost goes
Cache writes$1.30$0.80
Cache reads$0.40$0.40
Uncached input$0.20$0.20
Output$0.50$0.60
caching saves on the session with Claude Sonnet 5 (56%)
$3.10
caching saves on the session with Gemini 3.1 Pro Preview (64%)
$3.60

Is Gemini 3.1 Pro Preview cheaper than Claude Sonnet 5?

On input and cache hits the two are level: Claude Sonnet 5 and Gemini 3.1 Pro Preview both charge $2 per million input tokens and $0.20 per million cached tokens. Output costs $12 per million on Gemini and $10 on Sonnet 5. Cache writes are where Google comes out ahead: it publishes no write price, so written tokens cost ordinary input, while Anthropic charges $2.50 per million for a 5-minute write and $4 for a 1-hour write.

That split decides each workload. The example agentic session writes 400K tokens to the cache, and writes cost $1.30 on Sonnet 5 against $0.80 on Gemini. Gemini gives back $0.10 on output, $0.60 against $0.50, and the session ends at $2.00 against $2.40. A month of 110 sessions comes to $220.00 against $264.00. Without cache writes the order reverses: the one-off review is $0.40 on Sonnet 5 and $0.42 on Gemini, and the output-heavy generation $0.86 against $1.02.

What changes above 200K input tokens

Google prices long prompts separately for Gemini 3.1 Pro Preview. Once a prompt passes 200K input tokens, the rates become $4 input, $0.40 cached, and $18 output per million. Every request in the example session stays under that line, so none of the figures above include it.

Coding agents cross 200K more often than chat does. Loading a large module tree, a long test log, or a full dependency manifest can push a single request past it, and on those requests Gemini's input rate doubles and its output rate rises by half. If your sessions regularly run that large, compare them at the higher tier before reading anything into the session figure.

Google's explicit caching also carries a storage charge on Pro models, $4.50 per million tokens per hour, for as long as the cache lives. Implicit caching is free and on by default, but Google does not promise a hit on any given request.

Preview status, tool support, and thinking

Gemini 3.1 Pro Preview is Google's current Pro model, but it is still labeled a preview, and Google has announced Gemini 3.5 Pro without releasing it yet. GitHub Copilot retired Gemini 3.1 Pro Preview on September 1, 2026. It remains in Gemini CLI, where it is the Pro half of the default auto model, and in Cursor, OpenRouter, and OpenCode. Gemini API key users get the gemini-3.1-pro-preview-customtools endpoint at the same price.

Google describes the model as offering "advanced intelligence, complex problem-solving skills, and powerful agentic and vibe coding capabilities." Anthropic positions Sonnet 5 as its balance of speed and intelligence, close to Claude Opus 4.8 at lower prices. Both think by default: Gemini's thinking cannot be turned off and defaults to high, and Sonnet 5 runs adaptive thinking at high effort. With thinking on by default on both, output length drives more of the cost than the rate card suggests.

Sonnet 5 writes up to 128K tokens per response against Gemini's 65.5K, and both have context windows of roughly a million tokens, 1M and 1.05M.

Prompt caching

How each maker bills cached tokens

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

Google

Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.

Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.

On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.

Source: Google: Context caching

Your own numbers

See what Claude Sonnet 5 and Gemini 3.1 Pro Preview really cost you.

everyaitoken reads your Claude Code, Cursor, OpenCode, OpenRouter, and Gemini CLI history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Which is cheaper, Claude Sonnet 5 or Gemini 3.1 Pro Preview?

It depends on caching. On the example cached session Gemini costs $2.00 against $2.40. On uncached review and output-heavy generation Sonnet 5 costs less, and Gemini's rates rise past 200K input tokens.

Can I still pick Gemini 3.1 Pro Preview in GitHub Copilot?

No. Copilot dropped Gemini 3.1 Pro Preview on September 1, 2026, and it still offers Claude Sonnet 5.

How long a response can each model write?

Claude Sonnet 5 writes up to 128K tokens in one response and Gemini 3.1 Pro Preview up to 65.5K. For very long generated files, that difference can decide the choice on its own.

How do I compare my own Claude Code and Gemini CLI costs?

EveryToken reads both tools' local history on your Mac and prices each request at API rates, so you can see what each model costs you and what caching saved. The totals are API-equivalent estimates, not what a subscription charges.

  • Claude Fable 5.1 vs Claude Sonnet 5

    Claude Fable 5.1 lists at 5x the rates of Claude Sonnet 5, but cache hits cost $0.25 against $0.20. Why sessions cost 4.4x, and who Anthropic aims Fable 5.1 at.

  • Claude Fable 5.1 vs Gemini 3.1 Pro Preview

    A cached coding session costs 5.3x more on Claude Fable 5.1 than on Gemini 3.1 Pro Preview, mostly from cache writes. Output limits and preview status compared.

  • Claude Opus 4.8 vs Gemini 3.1 Pro Preview

    Claude Opus 4.8 costs 3x Gemini 3.1 Pro Preview on a cached coding session, $6.00 against $2.00. How long prompts, fast mode, and preview status shift that.

  • Claude Opus 5.5 vs Claude Sonnet 5

    Claude Opus 5.5 costs twice as much per token as Claude Sonnet 5, yet cache hits cost $0.20 on both. What that means for agentic coding sessions.

  • Claude Opus 5.5 vs Gemini 3.1 Pro Preview

    Gemini 3.1 Pro Preview costs less than half of Claude Opus 5.5 on a cached coding session, though both charge $0.20 per cache hit. Where the gap comes from.

  • Claude Sonnet 5 vs Claude Haiku 4.5

    Claude Haiku 4.5 costs half as much as Claude Sonnet 5, with a 200K context window and a retirement date ahead. When the cheaper model fits a coding workflow.