Model comparison
Claude Sonnet 5 vs Gemini 3.8 Flash: price, limits, caching
Gemini 3.8 Flash runs a cached coding session for $0.71 against $2.40 on Claude Sonnet 5, on introductory rates that end December 31, 2026.
· Prices as of September 28, 2026
Claude Sonnet 5
Anthropic · Released June 30, 2026
Anthropic's balance of speed and intelligence, and a drop-in upgrade for Claude Sonnet 4.6. Anthropic pitches it as close to Claude Opus 4.8 at a lower price.
Claude Sonnet 5 facts and comparisonsGemini 3.8 Flash
Google · Released September 2, 2026
Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.
Gemini 3.8 Flash facts and comparisons
The short answer
Gemini 3.8 Flash is much cheaper: its $0.75 input and $3.75 output rates are 63% below Claude Sonnet 5's, and the example agentic coding session costs $0.71 against $2.40, a 3.4x gap. Those are introductory rates that run through December 31, 2026, and Sonnet 5 can write up to 128K tokens in one response against 65.5K. Gemini 3.8 Flash suits low-cost work in Gemini CLI, and Sonnet 5 suits Claude Code and jobs that need long outputs.
Choose Claude Sonnet 5 if
- You need long single responses: Sonnet 5 writes up to 128K output tokens, about twice Gemini 3.8 Flash's 65.5K.
- You work in Claude Code, where the sonnet alias resolves to Sonnet 5 on the Anthropic API.
- You want rates with no scheduled change: Sonnet 5's launch price became its standard price on August 10, 2026.
- Anthropic's pitch matches your work: planning and using tools like browsers and terminals on its own.
Choose Gemini 3.8 Flash if
- You want a much lower cost per session: $0.71 against $2.40 for the example agentic session at introductory rates.
- You use Gemini CLI, where Gemini 3.8 Flash is the Flash half of the default auto model for Gemini API key and Vertex AI users.
- You want to start without paying: the Gemini API free tier covers its input, output, and caching.
- Your work is long-horizon software engineering or multi-file refactoring, which Google names among its strengths.
Side by side
Specs and prices
| Fact | Claude Sonnet 5 | Gemini 3.8 Flash |
|---|---|---|
| Maker | Anthropic | |
| API model id | claude-sonnet-5 | gemini-3.8-flash |
| Released | June 30, 2026 | September 2, 2026 |
| Status | Current | Current |
| Context window | 1M tokens | 1.05M tokens |
| Max output | 128K tokens | 65.5K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $2 | $0.75 |
| Cache hit, per 1M | $0.20 | $0.075 |
| Cache write, per 1M | $2.50 (5-minute), $4 (1-hour) | $0.75 (same as input) |
| Output, per 1M tokens | $10 | $3.75 |
| Runs in | Claude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot | Cursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker (Claude Sonnet 5: September 26, 2026; Gemini 3.8 Flash: September 28, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Claude Sonnet 5: The full 1M context window is billed at standard rates. The launch price became the standard price on August 10, 2026, and a planned increase was cancelled. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Claude Sonnet 5 | Gemini 3.8 Flash |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $2.40 | $0.71 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.40 | $0.15 |
| Output-heavy generation, 30K input, 80K output | $0.86 | $0.32 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $264.00 | $78.38 |
| Where the session’s cost goes | ||
| Cache writes | $1.30 | $0.30 |
| Cache reads | $0.40 | $0.15 |
| Uncached input | $0.20 | $0.08 |
| Output | $0.50 | $0.19 |
- caching saves on the session with Claude Sonnet 5 (56%)
- $3.10
- caching saves on the session with Gemini 3.8 Flash (66%)
- $1.35
Why the session gap is wider than the price gap
Gemini 3.8 Flash charges $0.75 per million input tokens, $0.075 per million cache hits, and $3.75 per million output tokens. Claude Sonnet 5 charges $2 for input, $0.20 for cache hits, and $10 for output. That is 2.7x on each, and it is the ratio you see on uncached work: the large one-off review costs $0.40 against $0.15, and the output-heavy generation $0.86 against $0.32.
The agentic session widens the gap to 3.4x. Google publishes no separate cache-write price, so written tokens cost ordinary input, $0.75 per million. Anthropic prices a 5-minute write at 1.25x input and a 1-hour write at 2x. In the session, writes cost $1.30 on Sonnet 5 against $0.30 on Gemini 3.8 Flash, and that $1.00 accounts for most of the $1.69 difference.
Across 110 sessions a month, that is $264.00 against $78.38, a $185.62 gap in API-equivalent cost. Caching saves 66% of the uncached cost on Gemini 3.8 Flash and 56% on Sonnet 5, again because Google adds no write premium.
Gemini 3.8 Flash's price after December 31, 2026
The rates above are introductory. From January 1, 2027, Google lists $1.50 per million input tokens, $0.15 per million cached tokens, and $7.50 per million output tokens, twice today's rates. They would still sit below every Sonnet 5 rate, but the per-token gap would narrow considerably.
Two caching details matter for any Gemini cost estimate. Implicit caching is on by default, but a hit is not assured, so a session can pay full input where this post assumes a read. Explicit caching gives a dependable discount but adds a storage charge while the cache lives, $0.50 to $1 per million tokens per hour on Flash models. The example session assumes the hits land and ignores storage.
Output limits, thinking levels, and tools
Both models accept roughly a million tokens of context: 1M on Sonnet 5 and 1.05M on Gemini 3.8 Flash. Output is where they part. Sonnet 5 can write up to 128K tokens in one response and Gemini 3.8 Flash 65.5K, which matters for long generated files or large refactors returned in one piece.
Google calls Gemini 3.8 Flash "our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows," and cites better robustness against prompt injection. Google sets its thinking level to medium by default, and Gemini CLI raises it to high. Sonnet 5 runs adaptive thinking at high effort by default. Higher thinking levels write more tokens, and the two tokenizers count text differently, so real sessions will not land exactly on these figures.
Both models are offered in Cursor, OpenRouter, OpenCode, and GitHub Copilot. Beyond that, Sonnet 5 runs in Claude Code and Gemini 3.8 Flash in Gemini CLI.
Prompt caching
How each maker bills cached tokens
Anthropic
Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.
A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.
Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.
In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.
Source: Anthropic: Prompt caching
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what Claude Sonnet 5 and Gemini 3.8 Flash really cost you.
everyaitoken reads your Claude Code, Cursor, OpenCode, OpenRouter, and Gemini CLI history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
How much cheaper is Gemini 3.8 Flash than Claude Sonnet 5?
Per token, 2.7x on input, output, and cache hits. On the example cached session, 3.4x: $0.71 against $2.40, because Google bills cache writes as ordinary input. Over a month of 110 sessions the difference is $185.62.
Does Gemini 3.8 Flash's price go up in 2027?
Its current rates are introductory through December 31, 2026. From January 1, 2027, Google lists $1.50 input, $0.15 cached, and $7.50 output per million tokens, which are still below Claude Sonnet 5's rates.
Can I use Gemini 3.8 Flash for free?
Yes, within the Gemini API free tier, which covers input, output, and caching for this model. Paid use is billed at the rates in this post, and subscription plans from either maker are priced differently from these API rates.
How can I see what each model costs in my own work?
EveryToken reads your local Claude Code and Gemini CLI history, prices every request at the maker's API rates, and shows the cache savings for each model. It runs on macOS for a $9 one-time price.
Sources
- Anthropic: Pricing
- Anthropic docs: Claude Sonnet 5
- Anthropic: Introducing Claude Sonnet 5
- Claude Code docs: Model configuration
- Cursor docs: Claude Sonnet 5
- OpenRouter: Claude Sonnet 5
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Google: Gemini API pricing
- Google docs: Gemini 3.8 Flash
- Google: Gemini models
- Google: Latest Gemini model
- Google DeepMind: Gemini 3.8 Flash model card
- Google: Gemini 3.8 Flash
- Gemini CLI source: model configuration
- Cursor docs: Gemini 3.8 Flash
- OpenRouter: Gemini 3.8 Flash
- Anthropic: Prompt caching
- Google: Context caching