Skip to content

Model comparison

Gemini 3.8 Flash vs Gemini 3.1 Pro Preview in Gemini CLI

Gemini CLI's auto model uses both Gemini 3.8 Flash and Gemini 3.1 Pro Preview. What each costs, why Flash rates double in 2027, and the Pro's preview status.

· Prices as of September 28, 2026

The short answer

Gemini 3.8 Flash costs about a third as much as Gemini 3.1 Pro Preview: the example agentic coding session is $0.71 against $2.00, and output-heavy work is 3.2x apart. Gemini CLI's default auto model uses both, as its Flash and Pro halves. Flash is on introductory rates that double on January 1, 2027, and the Pro Preview is still in preview, with Gemini 3.5 Pro announced but not yet released.

Choose Gemini 3.8 Flash if

  • You want the lower rates: $0.75 input and $3.75 output per million through the end of 2026.
  • Your work is long-horizon software engineering or complex multi-file refactoring, which Google names as 3.8 Flash strengths.
  • You'd like to prototype for free: the Gemini API free tier covers 3.8 Flash's input, output, and caching.
  • You use GitHub Copilot, which lists 3.8 Flash and retired 3.1 Pro Preview on September 1, 2026.

Choose Gemini 3.1 Pro Preview if

  • You want Google's current Pro model, positioned for deep reasoning and agentic coding, and can work with a preview.
  • Your agents rely on custom tools alongside bash, which the customtools endpoint is built to prioritize.

Side by side

Specs and prices

FactGemini 3.8 FlashGemini 3.1 Pro Preview
MakerGoogleGoogle
API model idgemini-3.8-flashgemini-3.1-pro-preview
ReleasedSeptember 2, 2026February 19, 2026
StatusCurrentPreview
Context window1.05M tokens1.05M tokens
Max output65.5K tokens65.5K tokens
Open weightsNoNo
Input, per 1M tokens$0.75$2
Cache hit, per 1M$0.075$0.20
Cache write, per 1M$0.75 (same as input)$2 (same as input)
Output, per 1M tokens$3.75$12
Runs inCursor, Gemini CLI, OpenCode, OpenRouter, and GitHub CopilotCursor, Gemini CLI, OpenCode, and OpenRouter

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens. Gemini 3.1 Pro Preview: Prompts over 200K input tokens cost $4 input, $0.40 cached, and $18 output per million tokens.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadGemini 3.8 FlashGemini 3.1 Pro Preview
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$0.71$2.00
Large one-off review, 150K input with no cache hits, 10K output$0.15$0.42
Output-heavy generation, 30K input, 80K output$0.32$1.02
A month of sessions, 110 sessions: 5 a day, 22 working days$78.38$220.00
Where the session’s cost goes
Cache writes$0.30$0.80
Cache reads$0.15$0.40
Uncached input$0.08$0.20
Output$0.19$0.60
caching saves on the session with Gemini 3.8 Flash (66%)
$1.35
caching saves on the session with Gemini 3.1 Pro Preview (64%)
$3.60

What Gemini 3.8 Flash and Gemini 3.1 Pro Preview cost

Gemini 3.8 Flash charges $0.75 per million input tokens, $0.075 per million cached, and $3.75 per million output. Gemini 3.1 Pro Preview charges $2, $0.20, and $12. Input and cache hits are 2.7x apart and output 3.2x, so the more a task writes, the wider the gap: the output-heavy generation costs $0.32 on Flash and $1.02 on the Pro Preview.

Google bills no separate cache write, so written tokens cost ordinary input, and a hit costs 10% of input on both. The example session costs $0.71 on Flash and $2.00 on the Pro Preview, 2.8x. Output is 30% of the Pro Preview session and 26% of the Flash session. At 110 sessions a month, the estimate is $78.38 against $220.00 at API rates.

Explicit caching, which gives a set discount on a cache you reference by name, adds storage for as long as the cache lives: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models. Implicit caching is on by default and applies the discount automatically when a request repeats a prefix Google has cached, but hits aren't assured.

Flash's introductory rates end after December 31, 2026

Google's current Gemini 3.8 Flash prices are introductory and run through December 31, 2026. From January 1, 2027, Google lists $1.50 input, $0.15 cached, and $7.50 output per million, 2x today's rates. At those prices Flash still lists below the Pro Preview's $2 input and $12 output, but the distance narrows, especially on input.

The Pro Preview has its own price step: prompts over 200K input tokens cost $4 input, $0.40 cached, and $18 output per million. The example session stays under 200K per request, so the tables use the standard tier. An agent that regularly sends larger prompts to the Pro Preview pays the higher tier on each of those requests.

The Gemini API free tier covers 3.8 Flash's input, output, and caching, which makes it a low-cost place to prototype an agent before moving to paid usage.

How Gemini CLI uses both models

Gemini CLI's default auto model pairs these two: 3.8 Flash is the Flash half for Gemini API key and Vertex AI users, and 3.1 Pro Preview is the Pro half. Gemini API key users get the gemini-3.1-pro-preview-customtools endpoint, which Google says is better at prioritizing custom tools alongside bash, at the same price. A session in auto mode mixes the two, so its cost lands between the two columns in the tables.

Thinking defaults differ. Flash's default thinking level is medium, and Gemini CLI sends high. The Pro Preview's thinking cannot be turned off, and its default level is high. Higher thinking levels tend to produce more output tokens, the line where the two models are 3.2x apart.

Google calls 3.8 Flash its most intelligent Flash model, "engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows," and credits it with multi-step planning and tool orchestration with fewer failed loops. It positions 3.1 Pro Preview, its current Pro model, for deep reasoning and agentic coding, and has announced Gemini 3.5 Pro without releasing it yet. GitHub Copilot retired 3.1 Pro Preview on September 1, 2026.

Prompt caching

How Google bills cached tokens

Google

Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.

Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.

On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.

Source: Google: Context caching

Your own numbers

See what Gemini 3.8 Flash and Gemini 3.1 Pro Preview really cost you.

everyaitoken reads your Gemini CLI, Cursor, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is Gemini 3.8 Flash cheaper than Gemini 3.1 Pro Preview?

Yes. At current rates the example agentic session costs $0.71 on 3.8 Flash and $2.00 on 3.1 Pro Preview. Flash's rates double on January 1, 2027, and it still lists below the Pro Preview after that.

Which model does Gemini CLI use by default?

Its default auto model uses both: Gemini 3.8 Flash as the Flash half for Gemini API key and Vertex AI users, and Gemini 3.1 Pro Preview as the Pro half.

Will Gemini 3.1 Pro Preview be replaced?

Google has announced Gemini 3.5 Pro but not released it. Until it arrives, 3.1 Pro Preview is Google's current Pro model, still in preview.

How can I see how Gemini CLI's auto mode splits my spend?

EveryToken reads your local Gemini CLI history on your Mac and prices each request at Google's API rates, by model, so you can see how much went to Flash and how much to Pro.

  • Gemini 3.1 Pro Preview vs Gemini 2.5 Pro

    Gemini 3.1 Pro Preview costs 60% more per input token than Gemini 2.5 Pro and is still a preview, while 2.5 Pro now limits new access. How to choose.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash

    Gemini 3.8 Flash costs about half of Gemini 3.5 Flash on introductory rates. What changes on January 1, 2027, and which one Gemini CLI picks for you.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash-Lite

    Google recommends Gemini 3.8 Flash and Gemini 3.5 Flash-Lite together for new projects. Where each fits, the 2.1x session gap, and what changes in 2027.

  • Gemini 3.8 Flash vs Gemini 3.7 Flash

    Gemini 3.8 Flash and Gemini 3.7 Flash cost the same, and both rates double on January 1, 2027. What differs is Gemini CLI, Cursor, and Google's own pitch.

  • Claude Fable 5.1 vs Gemini 3.1 Pro Preview

    A cached coding session costs 5.3x more on Claude Fable 5.1 than on Gemini 3.1 Pro Preview, mostly from cache writes. Output limits and preview status compared.

  • Claude Fable 5.1 vs Gemini 3.8 Flash

    A cached coding session costs 14.8x more on Claude Fable 5.1 than on Gemini 3.8 Flash. How Flash's introductory rates, output cap, and free tier factor in.