Model comparison
Gemini 3.8 Flash vs Gemini 3.7 Flash at identical rates
Gemini 3.8 Flash and Gemini 3.7 Flash cost the same, and both rates double on January 1, 2027. What differs is Gemini CLI, Cursor, and Google's own pitch.
· Prices as of September 28, 2026
Gemini 3.8 Flash
Google · Released September 2, 2026
Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.
Gemini 3.8 Flash facts and comparisonsGemini 3.7 Flash
Google · Released August 13, 2026 · Previous generation
Launched as Google's coding and agent Flash, now the previous generation that Google keeps fully supported.
Gemini 3.7 Flash facts and comparisons
The short answer
Gemini 3.8 Flash and Gemini 3.7 Flash share the same introductory rates, so the example agentic coding session costs $0.71 on either. The difference is support: Gemini 3.8 Flash is part of Gemini CLI's default auto model and is listed in Cursor, while Gemini 3.7 Flash is in neither. Google calls 3.8 Flash its most intelligent Flash model and 3.7 Flash its previous generation, so at equal prices the case for staying on 3.7 is a workflow already built around it.
Choose Gemini 3.8 Flash if
- You use Gemini CLI with a Gemini API key or Vertex AI, where 3.8 Flash is the Flash half of the default auto model.
- You pick models in Cursor, which lists Gemini 3.8 Flash but not Gemini 3.7 Flash.
- You want the Flash model Google engineered for long-horizon software engineering and autonomous agents.
- You want to start on the Gemini API free tier, which covers 3.8 Flash's input, output, and caching.
Choose Gemini 3.7 Flash if
- You have a workflow validated on Gemini 3.7 Flash through OpenRouter or OpenCode and want to keep it steady.
- You rely on the debugging and web UI gains Google claimed for 3.7 Flash and haven't checked them on 3.8 Flash yet.
Side by side
Specs and prices
| Fact | Gemini 3.8 Flash | Gemini 3.7 Flash |
|---|---|---|
| Maker | ||
| API model id | gemini-3.8-flash | gemini-3.7-flash |
| Released | September 2, 2026 | August 13, 2026 |
| Status | Current | Previous generation |
| Context window | 1.05M tokens | 1.05M tokens |
| Max output | 65.5K tokens | 65.5K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $0.75 | $0.75 |
| Cache hit, per 1M | $0.075 | $0.075 |
| Cache write, per 1M | $0.75 (same as input) | $0.75 (same as input) |
| Output, per 1M tokens | $3.75 | $3.75 |
| Runs in | Cursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot | OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens. Gemini 3.7 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Gemini 3.8 Flash | Gemini 3.7 Flash |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.71 | $0.71 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.15 | $0.15 |
| Output-heavy generation, 30K input, 80K output | $0.32 | $0.32 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $78.38 | $78.38 |
| Where the session’s cost goes | ||
| Cache writes | $0.30 | $0.30 |
| Cache reads | $0.15 | $0.15 |
| Uncached input | $0.08 | $0.08 |
| Output | $0.19 | $0.19 |
- caching saves on the session with Gemini 3.8 Flash (66%)
- $1.35
- caching saves on the session with Gemini 3.7 Flash (66%)
- $1.35
Do Gemini 3.8 Flash and Gemini 3.7 Flash cost the same?
Yes, rate for rate. Gemini 3.8 Flash and Gemini 3.7 Flash both charge $0.75 per million input tokens, $0.075 per million cached tokens, and $3.75 per million output tokens. Google bills no separate cache write, so written tokens cost ordinary input on both.
The example session costs $0.71 on either, the uncached review $0.15, and the output-heavy generation $0.32. A month of 110 sessions is $78.38 on each. Caching saves $1.35, or 66% of the uncached cost, and cache writes are the largest line at $0.30, 42% of the session.
Both prices are introductory. Google lists them through December 31, 2026, and from January 1, 2027, both move to $1.50 input, $0.15 cached, and $7.50 output per million tokens, 2x the current rates. The change hits both models alike, so it does not tip the choice, but it does change what a Flash-heavy workflow costs from 2027.
Which of the two runs in Gemini CLI and Cursor?
Gemini CLI's default auto model pairs a Pro model with a Flash model, and for Gemini API key and Vertex AI users the Flash half is Gemini 3.8 Flash. Gemini CLI sends it a high thinking level, above its medium default. Gemini 3.7 Flash is not in Gemini CLI's model list.
Cursor lists Gemini 3.8 Flash and not Gemini 3.7 Flash. Both are available through OpenRouter, OpenCode, and GitHub Copilot. The Gemini API free tier covers 3.8 Flash's input, output, and caching, and the sources here don't say the same for 3.7 Flash.
What Google says changed from 3.7 Flash to 3.8 Flash
Google launched Gemini 3.7 Flash on August 13, 2026, as its coding and agent Flash, citing higher first-pass code accuracy in debugging and issue resolution and more functional layouts in web UI generation. It now calls 3.7 Flash its previous-generation Flash for complex coding, agentic workflows, and reliable multi-step execution, and keeps it fully supported.
Gemini 3.8 Flash followed on September 2, 2026. Google calls it its most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows. It claims complex multi-file refactoring, deterministic tool execution, fewer failed loops in multi-step planning, and better robustness against prompt injection.
Neither model has larger limits: 1.05M of context and 65.5K of output on each. With prices equal, any cost difference between them comes from how many tokens each writes, and thinking level drives that. EveryToken reads your local Gemini CLI, OpenCode, and OpenRouter history and prices every request, so the effect of a switch shows up per model.
Prompt caching
How Google bills cached tokens
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what Gemini 3.8 Flash and Gemini 3.7 Flash really cost you.
everyaitoken reads your Gemini CLI, Cursor, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is Gemini 3.8 Flash more expensive than Gemini 3.7 Flash?
No. Both charge $0.75 input, $0.075 for cache hits, and $3.75 output per million tokens, on introductory rates through December 31, 2026.
What will Gemini 3.8 Flash cost in 2027?
From January 1, 2027, Google lists $1.50 input, $0.15 cached, and $7.50 output per million tokens. Gemini 3.7 Flash moves to the same rates.
Is Gemini 3.7 Flash deprecated?
No. Google calls it its previous-generation Flash and keeps it fully supported. It is not in Gemini CLI's model list or Cursor's, but it is available through OpenRouter, OpenCode, and GitHub Copilot.
Which Flash model does Gemini CLI use?
Gemini API key and Vertex AI users get Gemini 3.8 Flash as the Flash half of the default auto model. Accounts on Sign in with Google that haven't received the latest Flash get Gemini 3.5 Flash instead.
Sources
- Google: Gemini API pricing
- Google docs: Gemini 3.8 Flash
- Google: Gemini models
- Google: Latest Gemini model
- Google DeepMind: Gemini 3.8 Flash model card
- Google: Gemini 3.8 Flash
- Gemini CLI source: model configuration
- Cursor docs: Gemini 3.8 Flash
- OpenRouter: Gemini 3.8 Flash
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Google docs: Gemini 3.7 Flash
- Google: Introducing Gemini 3.7 Flash
- Cursor docs: Models
- OpenRouter: Gemini 3.7 Flash
- Google: Context caching