Skip to content

Model comparison

Claude Sonnet 5 vs Gemini 3.5 Flash: legacy Flash on price

Gemini 3.5 Flash, now Google's legacy Flash, costs $1.50 against $2.40 on Claude Sonnet 5 per cached coding session. Most of the gap is cache writes.

· Prices as of September 28, 2026

  • Claude Sonnet 5

    Anthropic · Released June 30, 2026

    Anthropic's balance of speed and intelligence, and a drop-in upgrade for Claude Sonnet 4.6. Anthropic pitches it as close to Claude Opus 4.8 at a lower price.

    Claude Sonnet 5 facts and comparisons
  • Gemini 3.5 Flash

    Google · Released May 19, 2026 · Previous generation

    Launched in May 2026 as Google's agent and coding Flash, now called its legacy Flash for routine, high-throughput work.

    Gemini 3.5 Flash facts and comparisons

The short answer

Gemini 3.5 Flash costs less, $1.50 against $2.40 for Claude Sonnet 5 on the example agentic coding session, but the gap nearly closes on output-heavy work because its $9 output rate sits close to Sonnet 5's $10. Google now calls Gemini 3.5 Flash its legacy Flash, and while newer Flash models are on introductory rates it costs more per token than they do. For new work the real comparison is Sonnet 5 against a newer Flash, and Gemini 3.5 Flash matters mostly where Gemini CLI still falls back to it.

Choose Claude Sonnet 5 if

  • Your work is output-heavy: at $10 against $9 per million output tokens, the two cost nearly the same, and Sonnet 5 writes up to 128K tokens per response against 65.5K.
  • You work in Claude Code or Cursor, and Cursor does not offer Gemini 3.5 Flash.
  • You want a current model: Anthropic calls Sonnet 5 its most agentic Sonnet yet, while Google has moved Gemini 3.5 Flash to legacy status.

Choose Gemini 3.5 Flash if

  • Your sessions lean on the cache: Google bills written tokens as ordinary input, which is most of the $0.90 session gap.
  • You use Gemini CLI with a Sign in with Google account that does not yet get the latest Flash, where Gemini 3.5 Flash is the Flash model.
  • Your workload is routine and high-throughput, the use Google now names for it.

Side by side

Specs and prices

FactClaude Sonnet 5Gemini 3.5 Flash
MakerAnthropicGoogle
API model idclaude-sonnet-5gemini-3.5-flash
ReleasedJune 30, 2026May 19, 2026
StatusCurrentPrevious generation
Context window1M tokens1.05M tokens
Max output128K tokens65.5K tokens
Open weightsNoNo
Input, per 1M tokens$2$1.50
Cache hit, per 1M$0.20$0.15
Cache write, per 1M$2.50 (5-minute), $4 (1-hour)$1.50 (same as input)
Output, per 1M tokens$10$9
Runs inClaude Code, Cursor, OpenCode, OpenRouter, and GitHub CopilotGemini CLI, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker (Claude Sonnet 5: September 26, 2026; Gemini 3.5 Flash: September 28, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Claude Sonnet 5: The full 1M context window is billed at standard rates. The launch price became the standard price on August 10, 2026, and a planned increase was cancelled.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadClaude Sonnet 5Gemini 3.5 Flash
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$2.40$1.50
Large one-off review, 150K input with no cache hits, 10K output$0.40$0.32
Output-heavy generation, 30K input, 80K output$0.86$0.77
A month of sessions, 110 sessions: 5 a day, 22 working days$264.00$165.00
Where the session’s cost goes
Cache writes$1.30$0.60
Cache reads$0.40$0.30
Uncached input$0.20$0.15
Output$0.50$0.45
caching saves on the session with Claude Sonnet 5 (56%)
$3.10
caching saves on the session with Gemini 3.5 Flash (64%)
$2.70

Where the $0.90 session gap comes from

Gemini 3.5 Flash charges $1.50 per million input tokens, $0.15 per million cache hits, and $9 per million output tokens. Claude Sonnet 5 lists $2 input, $0.20 cached, and $10 output. Input and cache hits are 25% cheaper on Gemini, but output is only 10% cheaper.

The session gap is larger than either of those because of cache writes. Google publishes no write price, so written tokens cost ordinary input, $1.50 per million. Anthropic charges $2.50 per million for a 5-minute write and $4 for a 1-hour write, and the example session splits its 400K written tokens evenly between the two. Writes come to $1.30 on Sonnet 5 against $0.60 on Gemini, a $0.70 difference.

Every other line is close. Reads differ by $0.10, and fresh input and output by $0.05 each. The session totals are $2.40 and $1.50, and at 110 sessions a month, $264.00 against $165.00.

On uncached work the two are close

Take the cache out and most of the gap disappears. The large one-off review, 150K tokens in and 10K out, costs $0.40 on Sonnet 5 and $0.32 on Gemini 3.5 Flash. The output-heavy generation costs $0.86 against $0.77, just $0.09 apart.

Output is 30% of Gemini 3.5 Flash's session cost and 21% of Sonnet 5's. Sonnet 5 runs adaptive thinking at high effort by default, and any thinking adds output tokens, which is the rate where these two are nearest. Their tokenizers differ, so the same prompt will not count identically on both, and small per-token gaps like these can move in either direction on real work.

Why a newer Flash may be the real alternative

Google launched Gemini 3.5 Flash in May 2026 as its agent and coding Flash. It now describes it as "our legacy Flash model, providing baseline speed and foundational performance for routine, high-throughput workloads." While Gemini 3.6 to 3.8 Flash are on introductory rates, Gemini 3.5 Flash costs more per token than they do.

That makes Gemini 3.8 Flash the more natural Google model to weigh against Sonnet 5 for new work. Gemini 3.5 Flash stays relevant in Gemini CLI, which uses it as the Flash model for Sign in with Google accounts that do not yet get the latest Flash, and on OpenRouter, OpenCode, and GitHub Copilot.

Sonnet 5 is a current model with a 1M context window and 128K of output. Gemini 3.5 Flash has a 1.05M window and 65.5K of output. Google's own claim for it is output four times faster than other frontier models by its measure. Anthropic's claim for Sonnet 5 is performance close to Claude Opus 4.8 at lower prices.

Prompt caching

How each maker bills cached tokens

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

Google

Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.

Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.

On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.

Source: Google: Context caching

Your own numbers

See what Claude Sonnet 5 and Gemini 3.5 Flash really cost you.

everyaitoken reads your Claude Code, Cursor, OpenCode, OpenRouter, and Gemini CLI history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is Gemini 3.5 Flash cheaper than Claude Sonnet 5?

Yes, on every rate, though by only 10% on output. The example cached session costs $1.50 against $2.40, and a month of 110 sessions $165.00 against $264.00.

Why does Google call Gemini 3.5 Flash legacy?

Google has released newer Flash models since May 2026 and now positions Gemini 3.5 Flash for routine, high-throughput work. While those newer models are on introductory rates, Gemini 3.5 Flash costs more per token than they do.

Does Gemini CLI still use Gemini 3.5 Flash?

Yes, as the Flash model for Sign in with Google accounts that do not yet get the latest Flash. Gemini API key and Vertex AI users get Gemini 3.8 Flash in the default auto model.

How can I see what my sessions cost on each?

EveryToken reads Claude Code and Gemini CLI logs on your Mac, applies each maker's API rates to every request, and shows cost and cache savings by model. Its totals are API-equivalent estimates, not what a Google or Anthropic subscription charges.

  • Claude Fable 5.1 vs Claude Sonnet 5

    Claude Fable 5.1 lists at 5x the rates of Claude Sonnet 5, but cache hits cost $0.25 against $0.20. Why sessions cost 4.4x, and who Anthropic aims Fable 5.1 at.

  • Claude Opus 5.5 vs Claude Sonnet 5

    Claude Opus 5.5 costs twice as much per token as Claude Sonnet 5, yet cache hits cost $0.20 on both. What that means for agentic coding sessions.

  • Claude Sonnet 5 vs Claude Haiku 4.5

    Claude Haiku 4.5 costs half as much as Claude Sonnet 5, with a 200K context window and a retirement date ahead. When the cheaper model fits a coding workflow.

  • Claude Sonnet 5 vs Claude Sonnet 4.6

    Claude Sonnet 5 costs 33% less per token than Claude Sonnet 4.6, but its tokenizer counts up to 1.35x as many tokens. What the upgrade saves in practice.

  • Claude Sonnet 5 vs Gemini 3.1 Pro Preview

    Claude Sonnet 5 and Gemini 3.1 Pro Preview both charge $2 input, but Gemini's rates rise past 200K tokens and it is still a preview. Session: $2.40 vs $2.00.

  • Claude Sonnet 5 vs Gemini 3.8 Flash

    Gemini 3.8 Flash runs a cached coding session for $0.71 against $2.40 on Claude Sonnet 5, on introductory rates that end December 31, 2026.