Model comparison
Gemini 3.8 Flash or Gemini 3.5 Flash-Lite for subagents?
Google recommends Gemini 3.8 Flash and Gemini 3.5 Flash-Lite together for new projects. Where each fits, the 2.1x session gap, and what changes in 2027.
· Prices as of September 28, 2026
Gemini 3.8 Flash
Google · Released September 2, 2026
Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.
Gemini 3.8 Flash facts and comparisonsGemini 3.5 Flash-Lite
Google · Released July 21, 2026
Google's current low-cost, low-latency tier for high-volume and subagent work, recommended with Gemini 3.8 Flash for new projects.
Gemini 3.5 Flash-Lite facts and comparisons
The short answer
Gemini 3.8 Flash costs 2.5x Gemini 3.5 Flash-Lite for input and 1.5x for output, so the example agentic coding session comes to $0.71 against $0.34. Google recommends the two together for new projects, with 3.8 Flash for long-horizon coding and agent work and 3.5 Flash-Lite for high-volume and subagent tasks. Gemini 3.8 Flash's introductory rates end on December 31, 2026, while Flash-Lite's rates carry no promotional note.
Choose Gemini 3.8 Flash if
- You want the model Google engineered for long-horizon software engineering, multi-file refactoring, and autonomous agents.
- Your Gemini CLI runs on a Gemini API key or Vertex AI, so its auto model already uses 3.8 Flash for Flash work.
- You work in Cursor or GitHub Copilot, which list Gemini 3.8 Flash but not Gemini 3.5 Flash-Lite.
Choose Gemini 3.5 Flash-Lite if
- You run subagent tasks or document parsing, the work Google names for Gemini 3.5 Flash-Lite.
- You want low latency, and Google cites about 350 output tokens per second for Flash-Lite.
- You want rates that are not introductory: $0.30 input and $2.50 output per million tokens.
- You use the flash-lite alias in Gemini CLI, which resolves to 3.5 Flash-Lite.
Side by side
Specs and prices
| Fact | Gemini 3.8 Flash | Gemini 3.5 Flash-Lite |
|---|---|---|
| Maker | ||
| API model id | gemini-3.8-flash | gemini-3.5-flash-lite |
| Released | September 2, 2026 | July 21, 2026 |
| Status | Current | Current |
| Context window | 1.05M tokens | 1.05M tokens |
| Max output | 65.5K tokens | 65.5K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $0.75 | $0.30 |
| Cache hit, per 1M | $0.075 | $0.03 |
| Cache write, per 1M | $0.75 (same as input) | $0.30 (same as input) |
| Output, per 1M tokens | $3.75 | $2.50 |
| Runs in | Cursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot | Gemini CLI, OpenCode, and OpenRouter |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Gemini 3.8 Flash | Gemini 3.5 Flash-Lite |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.71 | $0.34 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.15 | $0.07 |
| Output-heavy generation, 30K input, 80K output | $0.32 | $0.21 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $78.38 | $36.85 |
| Where the session’s cost goes | ||
| Cache writes | $0.30 | $0.12 |
| Cache reads | $0.15 | $0.06 |
| Uncached input | $0.08 | $0.03 |
| Output | $0.19 | $0.13 |
- caching saves on the session with Gemini 3.8 Flash (66%)
- $1.35
- caching saves on the session with Gemini 3.5 Flash-Lite (61%)
- $0.54
What Google recommends each model for
Google positions Gemini 3.8 Flash as its most intelligent Flash, aimed at long-horizon software engineering, autonomous agents, and enterprise workflows. It claims complex multi-file refactoring, deterministic tool execution, and multi-step planning with fewer failed loops.
Gemini 3.5 Flash-Lite is Google's low-cost, low-latency tier for high-throughput execution. Google names subagent tasks and document parsing as its uses, cites about 350 output tokens per second, and includes computer use as a built-in tool. Its default thinking level is minimal, against medium on 3.8 Flash.
Google recommends the two together for new projects. Read alongside the positioning, that suggests a main agent on 3.8 Flash handing narrow, repeated steps to Flash-Lite, rather than a choice of one over the other.
Why output-heavy work narrows the price gap
Input and caching sit 2.5x apart: $0.75 against $0.30 per million input tokens, and $0.075 against $0.03 for cache hits. Output is closer, $3.75 against $2.50, or 1.5x. Google bills no separate cache write on either, so written tokens cost ordinary input.
So the gap depends on the work. Input-heavy work shows most of the gap: $0.15 against $0.07 for the uncached review and $0.71 against $0.34 for the session, roughly 2.1x each. The output-heavy generation costs $0.32 against $0.21, only 1.5x. Output is 38% of Flash-Lite's session cost but 26% of 3.8 Flash's.
Over 110 sessions a month, 3.8 Flash comes to $78.38 and Flash-Lite to $36.85, a difference of $41.53. Caching saves $1.35 on 3.8 Flash, or 66%, and $0.54 on Flash-Lite, or 61%.
Introductory pricing, Gemini CLI, and where each runs
Gemini 3.8 Flash's rates are introductory through December 31, 2026. Starting January 1, 2027, its rates become $1.50 for input, $0.15 for cache hits, and $7.50 for output. Gemini 3.5 Flash-Lite's rates carry no introductory note, so the figures above describe the gap for the rest of 2026.
In Gemini CLI, 3.8 Flash is the Flash half of the default auto model for Gemini API key and Vertex AI users, and the CLI sends it a high thinking level. The flash-lite alias resolves to 3.5 Flash-Lite. Both run in OpenCode and on OpenRouter, while Cursor and GitHub Copilot list 3.8 Flash only.
On limits they tie, at a 1.05M context window and 65.5K of output. 3.8 Flash also has free-tier coverage on the Gemini API for input, output, and caching. EveryToken reads your local Gemini CLI history and prices each request at Google's rates, so you can see how much of a session ran on each model.
Prompt caching
How Google bills cached tokens
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what Gemini 3.8 Flash and Gemini 3.5 Flash-Lite really cost you.
everyaitoken reads your Gemini CLI, Cursor, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is Gemini 3.5 Flash-Lite cheaper than Gemini 3.8 Flash?
Yes. It charges $0.30 input and $2.50 output per million tokens against $0.75 and $3.75. The example session costs $0.34 against $0.71.
Which model does Gemini CLI use for flash-lite?
The flash-lite alias in Gemini CLI resolves to Gemini 3.5 Flash-Lite. The default auto model uses Gemini 3.8 Flash as its Flash half for Gemini API key and Vertex AI users.
Will Gemini 3.8 Flash get more expensive?
Yes, once its introductory rates end after December 31, 2026. The 2027 price list has it at $1.50 input, $0.15 cached, and $7.50 output per million tokens.
Is Gemini 3.5 Flash-Lite in Cursor?
No. It runs in Gemini CLI, OpenCode, and OpenRouter, but not in Cursor or GitHub Copilot.
Sources
- Google: Gemini API pricing
- Google docs: Gemini 3.8 Flash
- Google: Gemini models
- Google: Latest Gemini model
- Google DeepMind: Gemini 3.8 Flash model card
- Google: Gemini 3.8 Flash
- Gemini CLI source: model configuration
- Cursor docs: Gemini 3.8 Flash
- OpenRouter: Gemini 3.8 Flash
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Google docs: Gemini 3.5 Flash-Lite
- Google: Gemini 3.6 Flash and 3.5 Flash-Lite
- OpenRouter: Gemini 3.5 Flash-Lite
- Google: Context caching