Gemini 3.8 Flash: price, context window, and caching
Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.
Released September 2, 2026 · Prices as of September 28, 2026
In Google’s words
“Our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows.”
Facts
Specs and prices
| Fact | Gemini 3.8 Flash |
|---|---|
| Maker | |
| API model id | gemini-3.8-flash |
| Released | September 2, 2026 |
| Status | Current |
| Context window | 1.05M tokens |
| Max output | 65.5K tokens |
| Open weights | No |
| Input, per 1M tokens | $0.75 |
| Cache hit, per 1M | $0.075 |
| Cache write, per 1M | $0.75 (same as input) |
| Output, per 1M tokens | $3.75 |
| Runs in | Cursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by the maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.
Good to know
- The Flash half of Gemini CLI's default auto model for Gemini API key and Vertex AI users.
- The default thinking level is medium, and Gemini CLI sends high.
- The Gemini API free tier covers its input, output, and caching.
Cost
What typical work costs
Example token counts at Gemini 3.8 Flash’s published rates. On the agentic session, caching saves $1.35 against billing every token as ordinary input.
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.71 |
|---|---|
| Large one-off review, 150K input with no cache hits, 10K output | $0.15 |
| Output-heavy generation, 30K input, 80K output | $0.32 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $78.38 |
Prompt caching
How Google bills cached tokens
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Gemini 3.8 Flash compared
Claude Fable 5.1 vs Gemini 3.8 Flash
A cached coding session costs 14.8x more on Claude Fable 5.1 than on Gemini 3.8 Flash. How Flash's introductory rates, output cap, and free tier factor in.
Claude Haiku 4.5 vs Gemini 3.8 Flash
Gemini 3.8 Flash undercuts Claude Haiku 4.5 on introductory rates, $0.71 against $1.20 per cached session. From January 1, 2027, its rates top Haiku's.
Claude Sonnet 5 vs Gemini 3.8 Flash
Gemini 3.8 Flash runs a cached coding session for $0.71 against $2.40 on Claude Sonnet 5, on introductory rates that end December 31, 2026.
DeepSeek-V4.1-Flash vs Gemini 3.8 Flash
DeepSeek-V4.1-Flash costs $0.22 for a cached coding session that costs $0.71 on Gemini 3.8 Flash, and Gemini's introductory rates end December 31, 2026.
Gemini 3.8 Flash vs Gemini 3.5 Flash
Gemini 3.8 Flash costs about half of Gemini 3.5 Flash on introductory rates. What changes on January 1, 2027, and which one Gemini CLI picks for you.
Claude Opus 5.5 vs Gemini 3.8 Flash
Claude Opus 5.5 costs 6.2x as much as Gemini 3.8 Flash on a cached coding session at Flash's introductory rates. What changes in 2027, and where each fits.
Gemini 3.8 Flash vs Gemini 3.5 Flash-Lite
Google recommends Gemini 3.8 Flash and Gemini 3.5 Flash-Lite together for new projects. Where each fits, the 2.1x session gap, and what changes in 2027.
Gemini 3.8 Flash vs Gemini 3.1 Pro Preview
Gemini CLI's auto model uses both Gemini 3.8 Flash and Gemini 3.1 Pro Preview. What each costs, why Flash rates double in 2027, and the Pro's preview status.
Gemini 3.8 Flash vs Gemini 3.7 Flash
Gemini 3.8 Flash and Gemini 3.7 Flash cost the same, and both rates double on January 1, 2027. What differs is Gemini CLI, Cursor, and Google's own pitch.
GLM-5.3-Flash vs Gemini 3.8 Flash
GLM-5.3-Flash costs $0.16 for a cached coding session against $0.71 on Gemini 3.8 Flash, whose introductory rates end December 31, 2026. What else differs.
GPT-5.6 Terra vs Gemini 3.8 Flash
GPT-5.6 Terra costs 3.1x as much as Gemini 3.8 Flash on a cached coding session, $2.20 against $0.71, and its listed 2027 rates stay below Terra's.
GPT-6 Astra vs Gemini 3.8 Flash
GPT-6 Astra and Gemini 3.8 Flash launched a day apart at opposite ends of the price range, 14.8x apart on a cached coding session. What each is built for.
GPT-6 Luna vs Gemini 3.8 Flash
Gemini 3.8 Flash costs 7.5x as much as GPT-6 Luna per token, and its introductory rates end December 31, 2026. A month of sessions: $11.55 vs $78.38.
GPT-6 Sol vs Gemini 3.8 Flash
Gemini 3.8 Flash costs a third of GPT-6 Sol on a cached coding session, at introductory rates that end December 31, 2026. How the gap changes after that.
Grok Build 0.1 vs Gemini 3.8 Flash
Gemini 3.8 Flash costs $0.71 per cached coding session against $1.00 on Grok Build 0.1, until its introductory rates end on December 31, 2026.
MiniMax M3 vs Gemini 3.8 Flash
MiniMax M3 costs $0.33 per cached coding session against $0.71 on Gemini 3.8 Flash, whose introductory rates end on December 31, 2026.
Mistral Medium 3.5 vs Gemini 3.8 Flash
Gemini 3.8 Flash costs half as much as Mistral Medium 3.5 on every rate until December 31, 2026. From January 1, 2027, their list prices match to the cent.
Your own numbers
See what Gemini 3.8 Flash really costs you.
everyaitoken reads your Gemini CLI, Cursor, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
Sources
- Google: Gemini API pricing
- Google docs: Gemini 3.8 Flash
- Google: Gemini models
- Google: Latest Gemini model
- Google DeepMind: Gemini 3.8 Flash model card
- Google: Gemini 3.8 Flash
- Gemini CLI source: model configuration
- Cursor docs: Gemini 3.8 Flash
- OpenRouter: Gemini 3.8 Flash
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Google: Context caching