Model comparison
Gemini 3.5 Flash to Gemini 3.8 Flash: far cheaper until 2027
Gemini 3.8 Flash costs about half of Gemini 3.5 Flash on introductory rates. What changes on January 1, 2027, and which one Gemini CLI picks for you.
· Prices as of September 28, 2026
Gemini 3.8 Flash
Google · Released September 2, 2026
Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.
Gemini 3.8 Flash facts and comparisonsGemini 3.5 Flash
Google · Released May 19, 2026 · Previous generation
Launched in May 2026 as Google's agent and coding Flash, now called its legacy Flash for routine, high-throughput work.
Gemini 3.5 Flash facts and comparisons
The short answer
Gemini 3.8 Flash is the newer model and, on its introductory rates, the cheaper one: the example agentic coding session costs $0.71 against $1.50 on Gemini 3.5 Flash. From January 1, 2027, 3.8 Flash moves to $1.50 input, $0.15 cached, and $7.50 output per million tokens, matching 3.5 Flash on input and caching and staying below its $9 output rate. Google now calls 3.5 Flash its legacy Flash, though Gemini CLI still uses it for some Sign in with Google accounts.
Choose Gemini 3.8 Flash if
- You want the lower price now: $0.75 input and $3.75 output per million tokens until December 31, 2026.
- You want the model Google calls its most intelligent Flash, built for long-horizon software engineering and autonomous agents.
- You use Cursor, which lists Gemini 3.8 Flash but not Gemini 3.5 Flash.
- You are prototyping on the Gemini API free tier, where 3.8 Flash is included.
Choose Gemini 3.5 Flash if
- You sign in to Gemini CLI with a Google account that doesn't yet get the latest Flash, where 3.5 Flash is the Flash model.
- You run routine, high-throughput workloads, the role Google now gives its legacy Flash.
- You weigh output speed, which Google measured at 3.5 Flash's launch as four times faster than other frontier models; the sources here give no speed figure for 3.8 Flash.
Side by side
Specs and prices
| Fact | Gemini 3.8 Flash | Gemini 3.5 Flash |
|---|---|---|
| Maker | ||
| API model id | gemini-3.8-flash | gemini-3.5-flash |
| Released | September 2, 2026 | May 19, 2026 |
| Status | Current | Previous generation |
| Context window | 1.05M tokens | 1.05M tokens |
| Max output | 65.5K tokens | 65.5K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $0.75 | $1.50 |
| Cache hit, per 1M | $0.075 | $0.15 |
| Cache write, per 1M | $0.75 (same as input) | $1.50 (same as input) |
| Output, per 1M tokens | $3.75 | $9 |
| Runs in | Cursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot | Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Gemini 3.8 Flash | Gemini 3.5 Flash |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.71 | $1.50 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.15 | $0.32 |
| Output-heavy generation, 30K input, 80K output | $0.32 | $0.77 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $78.38 | $165.00 |
| Where the session’s cost goes | ||
| Cache writes | $0.30 | $0.60 |
| Cache reads | $0.15 | $0.30 |
| Uncached input | $0.08 | $0.15 |
| Output | $0.19 | $0.45 |
- caching saves on the session with Gemini 3.8 Flash (66%)
- $1.35
- caching saves on the session with Gemini 3.5 Flash (64%)
- $2.70
Why the newer Gemini Flash costs less
Gemini 3.5 Flash lists $1.50 per million input tokens and $9 per million output tokens, with cache hits at $0.15. Gemini 3.8 Flash, on introductory rates, lists $0.75 and $3.75, with cache hits at $0.075. Input and caching are 2x apart, and output is 2.4x apart.
The introductory rates explain the inversion. While Gemini 3.6 to 3.8 Flash are on introductory rates, 3.5 Flash costs more per token than they do. On the example workloads, 3.8 Flash costs $0.71 for the session against $1.50, $0.15 for the uncached review against $0.32, and $0.32 for the output-heavy generation against $0.77.
Across 110 sessions a month that is $78.38 against $165.00, a difference of $86.62. Caching saves 66% of the uncached session on 3.8 Flash and 64% on 3.5 Flash, and neither model carries a cache-write surcharge.
What changes on January 1, 2027
Gemini 3.8 Flash's introductory rates run through December 31, 2026. On January 1, 2027, its price becomes $1.50 per million input tokens, $0.15 per million cached tokens, and $7.50 per million output tokens.
Those match Gemini 3.5 Flash's current rates for input and cache hits exactly. Output stays lower, $7.50 against $9. After the promotion, then, the price case for 3.8 Flash narrows to output, and the choice rests more on what each model is for.
Gemini 3.5 Flash's own rates carry no promotional note in the sources here, so this comparison assumes they stay where they are.
Gemini CLI still uses both Flash models
Gemini CLI picks a Flash model by account type. For Gemini API key and Vertex AI users, the Flash half of the default auto model is 3.8 Flash, sent a high thinking level. For Sign in with Google accounts that don't yet get the latest Flash, Gemini CLI uses 3.5 Flash.
Google launched 3.5 Flash in May 2026 as its agent and coding Flash and now calls it its legacy Flash for routine, high-throughput workloads. At launch Google said its output ran four times faster than other frontier models by its own measure, and it names sub-agent deployment and long-horizon tasks at scale among its uses.
Gemini 3.8 Flash arrived on September 2, 2026. Google claims long-horizon software engineering, complex multi-file refactoring, fewer failed loops in multi-step planning, and better robustness against prompt injection. Each accepts a 1.05M context window and writes up to 65.5K tokens. If your Gemini CLI sessions land on different Flash models by account, EveryToken prices each local request at the rate of the model that actually ran.
Prompt caching
How Google bills cached tokens
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what Gemini 3.8 Flash and Gemini 3.5 Flash really cost you.
everyaitoken reads your Gemini CLI, Cursor, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is Gemini 3.8 Flash cheaper than Gemini 3.5 Flash?
Yes, while its introductory rates last. It charges $0.75 input and $3.75 output per million tokens against $1.50 and $9 for 3.5 Flash, through December 31, 2026.
Will Gemini 3.8 Flash cost the same as Gemini 3.5 Flash in 2027?
For input and cache hits, yes: from January 1, 2027, 3.8 Flash charges $1.50 input and $0.15 cached. Its output rate becomes $7.50 per million tokens, still below the $9 of 3.5 Flash.
Why does Gemini CLI use Gemini 3.5 Flash for my account?
Gemini CLI uses 3.5 Flash as the Flash model for Sign in with Google accounts that don't yet get the latest Flash. Gemini API key and Vertex AI users get 3.8 Flash in the default auto model.
Is Gemini 3.5 Flash deprecated?
The sources here list no shutdown date. Google calls it its legacy Flash, for routine, high-throughput workloads.
Sources
- Google: Gemini API pricing
- Google docs: Gemini 3.8 Flash
- Google: Gemini models
- Google: Latest Gemini model
- Google DeepMind: Gemini 3.8 Flash model card
- Google: Gemini 3.8 Flash
- Gemini CLI source: model configuration
- Cursor docs: Gemini 3.8 Flash
- OpenRouter: Gemini 3.8 Flash
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Google docs: Gemini 3.5 Flash
- Google: Gemini 3.5
- OpenRouter: Gemini 3.5 Flash
- Google: Context caching