Skip to content

Model comparison

Claude Sonnet 5 vs Gemini 3.8 Flash: price, limits, caching

Gemini 3.8 Flash runs a cached coding session for $0.71 against $2.40 on Claude Sonnet 5, on introductory rates that end December 31, 2026.

· Prices as of September 28, 2026

  • Claude Sonnet 5

    Anthropic · Released June 30, 2026

    Anthropic's balance of speed and intelligence, and a drop-in upgrade for Claude Sonnet 4.6. Anthropic pitches it as close to Claude Opus 4.8 at a lower price.

    Claude Sonnet 5 facts and comparisons
  • Gemini 3.8 Flash

    Google · Released September 2, 2026

    Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.

    Gemini 3.8 Flash facts and comparisons

The short answer

Gemini 3.8 Flash is much cheaper: its $0.75 input and $3.75 output rates are 63% below Claude Sonnet 5's, and the example agentic coding session costs $0.71 against $2.40, a 3.4x gap. Those are introductory rates that run through December 31, 2026, and Sonnet 5 can write up to 128K tokens in one response against 65.5K. Gemini 3.8 Flash suits low-cost work in Gemini CLI, and Sonnet 5 suits Claude Code and jobs that need long outputs.

Choose Claude Sonnet 5 if

  • You need long single responses: Sonnet 5 writes up to 128K output tokens, about twice Gemini 3.8 Flash's 65.5K.
  • You work in Claude Code, where the sonnet alias resolves to Sonnet 5 on the Anthropic API.
  • You want rates with no scheduled change: Sonnet 5's launch price became its standard price on August 10, 2026.
  • Anthropic's pitch matches your work: planning and using tools like browsers and terminals on its own.

Choose Gemini 3.8 Flash if

  • You want a much lower cost per session: $0.71 against $2.40 for the example agentic session at introductory rates.
  • You use Gemini CLI, where Gemini 3.8 Flash is the Flash half of the default auto model for Gemini API key and Vertex AI users.
  • You want to start without paying: the Gemini API free tier covers its input, output, and caching.
  • Your work is long-horizon software engineering or multi-file refactoring, which Google names among its strengths.

Side by side

Specs and prices

FactClaude Sonnet 5Gemini 3.8 Flash
MakerAnthropicGoogle
API model idclaude-sonnet-5gemini-3.8-flash
ReleasedJune 30, 2026September 2, 2026
StatusCurrentCurrent
Context window1M tokens1.05M tokens
Max output128K tokens65.5K tokens
Open weightsNoNo
Input, per 1M tokens$2$0.75
Cache hit, per 1M$0.20$0.075
Cache write, per 1M$2.50 (5-minute), $4 (1-hour)$0.75 (same as input)
Output, per 1M tokens$10$3.75
Runs inClaude Code, Cursor, OpenCode, OpenRouter, and GitHub CopilotCursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker (Claude Sonnet 5: September 26, 2026; Gemini 3.8 Flash: September 28, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Claude Sonnet 5: The full 1M context window is billed at standard rates. The launch price became the standard price on August 10, 2026, and a planned increase was cancelled. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadClaude Sonnet 5Gemini 3.8 Flash
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$2.40$0.71
Large one-off review, 150K input with no cache hits, 10K output$0.40$0.15
Output-heavy generation, 30K input, 80K output$0.86$0.32
A month of sessions, 110 sessions: 5 a day, 22 working days$264.00$78.38
Where the session’s cost goes
Cache writes$1.30$0.30
Cache reads$0.40$0.15
Uncached input$0.20$0.08
Output$0.50$0.19
caching saves on the session with Claude Sonnet 5 (56%)
$3.10
caching saves on the session with Gemini 3.8 Flash (66%)
$1.35

Why the session gap is wider than the price gap

Gemini 3.8 Flash charges $0.75 per million input tokens, $0.075 per million cache hits, and $3.75 per million output tokens. Claude Sonnet 5 charges $2 for input, $0.20 for cache hits, and $10 for output. That is 2.7x on each, and it is the ratio you see on uncached work: the large one-off review costs $0.40 against $0.15, and the output-heavy generation $0.86 against $0.32.

The agentic session widens the gap to 3.4x. Google publishes no separate cache-write price, so written tokens cost ordinary input, $0.75 per million. Anthropic prices a 5-minute write at 1.25x input and a 1-hour write at 2x. In the session, writes cost $1.30 on Sonnet 5 against $0.30 on Gemini 3.8 Flash, and that $1.00 accounts for most of the $1.69 difference.

Across 110 sessions a month, that is $264.00 against $78.38, a $185.62 gap in API-equivalent cost. Caching saves 66% of the uncached cost on Gemini 3.8 Flash and 56% on Sonnet 5, again because Google adds no write premium.

Gemini 3.8 Flash's price after December 31, 2026

The rates above are introductory. From January 1, 2027, Google lists $1.50 per million input tokens, $0.15 per million cached tokens, and $7.50 per million output tokens, twice today's rates. They would still sit below every Sonnet 5 rate, but the per-token gap would narrow considerably.

Two caching details matter for any Gemini cost estimate. Implicit caching is on by default, but a hit is not assured, so a session can pay full input where this post assumes a read. Explicit caching gives a dependable discount but adds a storage charge while the cache lives, $0.50 to $1 per million tokens per hour on Flash models. The example session assumes the hits land and ignores storage.

Output limits, thinking levels, and tools

Both models accept roughly a million tokens of context: 1M on Sonnet 5 and 1.05M on Gemini 3.8 Flash. Output is where they part. Sonnet 5 can write up to 128K tokens in one response and Gemini 3.8 Flash 65.5K, which matters for long generated files or large refactors returned in one piece.

Google calls Gemini 3.8 Flash "our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows," and cites better robustness against prompt injection. Google sets its thinking level to medium by default, and Gemini CLI raises it to high. Sonnet 5 runs adaptive thinking at high effort by default. Higher thinking levels write more tokens, and the two tokenizers count text differently, so real sessions will not land exactly on these figures.

Both models are offered in Cursor, OpenRouter, OpenCode, and GitHub Copilot. Beyond that, Sonnet 5 runs in Claude Code and Gemini 3.8 Flash in Gemini CLI.

Prompt caching

How each maker bills cached tokens

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

Google

Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.

Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.

On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.

Source: Google: Context caching

Your own numbers

See what Claude Sonnet 5 and Gemini 3.8 Flash really cost you.

everyaitoken reads your Claude Code, Cursor, OpenCode, OpenRouter, and Gemini CLI history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

How much cheaper is Gemini 3.8 Flash than Claude Sonnet 5?

Per token, 2.7x on input, output, and cache hits. On the example cached session, 3.4x: $0.71 against $2.40, because Google bills cache writes as ordinary input. Over a month of 110 sessions the difference is $185.62.

Does Gemini 3.8 Flash's price go up in 2027?

Its current rates are introductory through December 31, 2026. From January 1, 2027, Google lists $1.50 input, $0.15 cached, and $7.50 output per million tokens, which are still below Claude Sonnet 5's rates.

Can I use Gemini 3.8 Flash for free?

Yes, within the Gemini API free tier, which covers input, output, and caching for this model. Paid use is billed at the rates in this post, and subscription plans from either maker are priced differently from these API rates.

How can I see what each model costs in my own work?

EveryToken reads your local Claude Code and Gemini CLI history, prices every request at the maker's API rates, and shows the cache savings for each model. It runs on macOS for a $9 one-time price.

  • Claude Fable 5.1 vs Claude Sonnet 5

    Claude Fable 5.1 lists at 5x the rates of Claude Sonnet 5, but cache hits cost $0.25 against $0.20. Why sessions cost 4.4x, and who Anthropic aims Fable 5.1 at.

  • Claude Fable 5.1 vs Gemini 3.8 Flash

    A cached coding session costs 14.8x more on Claude Fable 5.1 than on Gemini 3.8 Flash. How Flash's introductory rates, output cap, and free tier factor in.

  • Claude Haiku 4.5 vs Gemini 3.8 Flash

    Gemini 3.8 Flash undercuts Claude Haiku 4.5 on introductory rates, $0.71 against $1.20 per cached session. From January 1, 2027, its rates top Haiku's.

  • Claude Opus 5.5 vs Claude Sonnet 5

    Claude Opus 5.5 costs twice as much per token as Claude Sonnet 5, yet cache hits cost $0.20 on both. What that means for agentic coding sessions.

  • Claude Sonnet 5 vs Claude Haiku 4.5

    Claude Haiku 4.5 costs half as much as Claude Sonnet 5, with a 200K context window and a retirement date ahead. When the cheaper model fits a coding workflow.

  • Claude Sonnet 5 vs Claude Sonnet 4.6

    Claude Sonnet 5 costs 33% less per token than Claude Sonnet 4.6, but its tokenizer counts up to 1.35x as many tokens. What the upgrade saves in practice.