Model comparison
Gemini 3.8 Flash vs Gemini 3.1 Pro Preview in Gemini CLI
Gemini CLI's auto model uses both Gemini 3.8 Flash and Gemini 3.1 Pro Preview. What each costs, why Flash rates double in 2027, and the Pro's preview status.
· Prices as of September 28, 2026
Gemini 3.8 Flash
Google · Released September 2, 2026
Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.
Gemini 3.8 Flash facts and comparisonsGemini 3.1 Pro Preview
Google · Released February 19, 2026 · Preview
Google's only current Pro model, still in preview, positioned for deep reasoning and agentic coding.
Gemini 3.1 Pro Preview facts and comparisons
The short answer
Gemini 3.8 Flash costs about a third as much as Gemini 3.1 Pro Preview: the example agentic coding session is $0.71 against $2.00, and output-heavy work is 3.2x apart. Gemini CLI's default auto model uses both, as its Flash and Pro halves. Flash is on introductory rates that double on January 1, 2027, and the Pro Preview is still in preview, with Gemini 3.5 Pro announced but not yet released.
Choose Gemini 3.8 Flash if
- You want the lower rates: $0.75 input and $3.75 output per million through the end of 2026.
- Your work is long-horizon software engineering or complex multi-file refactoring, which Google names as 3.8 Flash strengths.
- You'd like to prototype for free: the Gemini API free tier covers 3.8 Flash's input, output, and caching.
- You use GitHub Copilot, which lists 3.8 Flash and retired 3.1 Pro Preview on September 1, 2026.
Choose Gemini 3.1 Pro Preview if
- You want Google's current Pro model, positioned for deep reasoning and agentic coding, and can work with a preview.
- Your agents rely on custom tools alongside bash, which the customtools endpoint is built to prioritize.
Side by side
Specs and prices
| Fact | Gemini 3.8 Flash | Gemini 3.1 Pro Preview |
|---|---|---|
| Maker | ||
| API model id | gemini-3.8-flash | gemini-3.1-pro-preview |
| Released | September 2, 2026 | February 19, 2026 |
| Status | Current | Preview |
| Context window | 1.05M tokens | 1.05M tokens |
| Max output | 65.5K tokens | 65.5K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $0.75 | $2 |
| Cache hit, per 1M | $0.075 | $0.20 |
| Cache write, per 1M | $0.75 (same as input) | $2 (same as input) |
| Output, per 1M tokens | $3.75 | $12 |
| Runs in | Cursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot | Cursor, Gemini CLI, OpenCode, and OpenRouter |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens. Gemini 3.1 Pro Preview: Prompts over 200K input tokens cost $4 input, $0.40 cached, and $18 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Gemini 3.8 Flash | Gemini 3.1 Pro Preview |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.71 | $2.00 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.15 | $0.42 |
| Output-heavy generation, 30K input, 80K output | $0.32 | $1.02 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $78.38 | $220.00 |
| Where the session’s cost goes | ||
| Cache writes | $0.30 | $0.80 |
| Cache reads | $0.15 | $0.40 |
| Uncached input | $0.08 | $0.20 |
| Output | $0.19 | $0.60 |
- caching saves on the session with Gemini 3.8 Flash (66%)
- $1.35
- caching saves on the session with Gemini 3.1 Pro Preview (64%)
- $3.60
What Gemini 3.8 Flash and Gemini 3.1 Pro Preview cost
Gemini 3.8 Flash charges $0.75 per million input tokens, $0.075 per million cached, and $3.75 per million output. Gemini 3.1 Pro Preview charges $2, $0.20, and $12. Input and cache hits are 2.7x apart and output 3.2x, so the more a task writes, the wider the gap: the output-heavy generation costs $0.32 on Flash and $1.02 on the Pro Preview.
Google bills no separate cache write, so written tokens cost ordinary input, and a hit costs 10% of input on both. The example session costs $0.71 on Flash and $2.00 on the Pro Preview, 2.8x. Output is 30% of the Pro Preview session and 26% of the Flash session. At 110 sessions a month, the estimate is $78.38 against $220.00 at API rates.
Explicit caching, which gives a set discount on a cache you reference by name, adds storage for as long as the cache lives: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models. Implicit caching is on by default and applies the discount automatically when a request repeats a prefix Google has cached, but hits aren't assured.
Flash's introductory rates end after December 31, 2026
Google's current Gemini 3.8 Flash prices are introductory and run through December 31, 2026. From January 1, 2027, Google lists $1.50 input, $0.15 cached, and $7.50 output per million, 2x today's rates. At those prices Flash still lists below the Pro Preview's $2 input and $12 output, but the distance narrows, especially on input.
The Pro Preview has its own price step: prompts over 200K input tokens cost $4 input, $0.40 cached, and $18 output per million. The example session stays under 200K per request, so the tables use the standard tier. An agent that regularly sends larger prompts to the Pro Preview pays the higher tier on each of those requests.
The Gemini API free tier covers 3.8 Flash's input, output, and caching, which makes it a low-cost place to prototype an agent before moving to paid usage.
How Gemini CLI uses both models
Gemini CLI's default auto model pairs these two: 3.8 Flash is the Flash half for Gemini API key and Vertex AI users, and 3.1 Pro Preview is the Pro half. Gemini API key users get the gemini-3.1-pro-preview-customtools endpoint, which Google says is better at prioritizing custom tools alongside bash, at the same price. A session in auto mode mixes the two, so its cost lands between the two columns in the tables.
Thinking defaults differ. Flash's default thinking level is medium, and Gemini CLI sends high. The Pro Preview's thinking cannot be turned off, and its default level is high. Higher thinking levels tend to produce more output tokens, the line where the two models are 3.2x apart.
Google calls 3.8 Flash its most intelligent Flash model, "engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows," and credits it with multi-step planning and tool orchestration with fewer failed loops. It positions 3.1 Pro Preview, its current Pro model, for deep reasoning and agentic coding, and has announced Gemini 3.5 Pro without releasing it yet. GitHub Copilot retired 3.1 Pro Preview on September 1, 2026.
Prompt caching
How Google bills cached tokens
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what Gemini 3.8 Flash and Gemini 3.1 Pro Preview really cost you.
everyaitoken reads your Gemini CLI, Cursor, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is Gemini 3.8 Flash cheaper than Gemini 3.1 Pro Preview?
Yes. At current rates the example agentic session costs $0.71 on 3.8 Flash and $2.00 on 3.1 Pro Preview. Flash's rates double on January 1, 2027, and it still lists below the Pro Preview after that.
Which model does Gemini CLI use by default?
Its default auto model uses both: Gemini 3.8 Flash as the Flash half for Gemini API key and Vertex AI users, and Gemini 3.1 Pro Preview as the Pro half.
Will Gemini 3.1 Pro Preview be replaced?
Google has announced Gemini 3.5 Pro but not released it. Until it arrives, 3.1 Pro Preview is Google's current Pro model, still in preview.
How can I see how Gemini CLI's auto mode splits my spend?
EveryToken reads your local Gemini CLI history on your Mac and prices each request at Google's API rates, by model, so you can see how much went to Flash and how much to Pro.
Sources
- Google: Gemini API pricing
- Google docs: Gemini 3.8 Flash
- Google: Gemini models
- Google: Latest Gemini model
- Google DeepMind: Gemini 3.8 Flash model card
- Google: Gemini 3.8 Flash
- Gemini CLI source: model configuration
- Cursor docs: Gemini 3.8 Flash
- OpenRouter: Gemini 3.8 Flash
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Google docs: Gemini 3.1 Pro Preview
- Google: Gemini 3.1 Pro
- Google DeepMind: Gemini
- Cursor docs: Gemini 3.1 Pro
- OpenRouter: Gemini 3.1 Pro Preview
- Google: Context caching