Skip to content

Model comparison

Gemini 3.8 Flash or Gemini 3.5 Flash-Lite for subagents?

Google recommends Gemini 3.8 Flash and Gemini 3.5 Flash-Lite together for new projects. Where each fits, the 2.1x session gap, and what changes in 2027.

· Prices as of September 28, 2026

  • Gemini 3.8 Flash

    Google · Released September 2, 2026

    Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.

    Gemini 3.8 Flash facts and comparisons
  • Gemini 3.5 Flash-Lite

    Google · Released July 21, 2026

    Google's current low-cost, low-latency tier for high-volume and subagent work, recommended with Gemini 3.8 Flash for new projects.

    Gemini 3.5 Flash-Lite facts and comparisons

The short answer

Gemini 3.8 Flash costs 2.5x Gemini 3.5 Flash-Lite for input and 1.5x for output, so the example agentic coding session comes to $0.71 against $0.34. Google recommends the two together for new projects, with 3.8 Flash for long-horizon coding and agent work and 3.5 Flash-Lite for high-volume and subagent tasks. Gemini 3.8 Flash's introductory rates end on December 31, 2026, while Flash-Lite's rates carry no promotional note.

Choose Gemini 3.8 Flash if

  • You want the model Google engineered for long-horizon software engineering, multi-file refactoring, and autonomous agents.
  • Your Gemini CLI runs on a Gemini API key or Vertex AI, so its auto model already uses 3.8 Flash for Flash work.
  • You work in Cursor or GitHub Copilot, which list Gemini 3.8 Flash but not Gemini 3.5 Flash-Lite.

Choose Gemini 3.5 Flash-Lite if

  • You run subagent tasks or document parsing, the work Google names for Gemini 3.5 Flash-Lite.
  • You want low latency, and Google cites about 350 output tokens per second for Flash-Lite.
  • You want rates that are not introductory: $0.30 input and $2.50 output per million tokens.
  • You use the flash-lite alias in Gemini CLI, which resolves to 3.5 Flash-Lite.

Side by side

Specs and prices

FactGemini 3.8 FlashGemini 3.5 Flash-Lite
MakerGoogleGoogle
API model idgemini-3.8-flashgemini-3.5-flash-lite
ReleasedSeptember 2, 2026July 21, 2026
StatusCurrentCurrent
Context window1.05M tokens1.05M tokens
Max output65.5K tokens65.5K tokens
Open weightsNoNo
Input, per 1M tokens$0.75$0.30
Cache hit, per 1M$0.075$0.03
Cache write, per 1M$0.75 (same as input)$0.30 (same as input)
Output, per 1M tokens$3.75$2.50
Runs inCursor, Gemini CLI, OpenCode, OpenRouter, and GitHub CopilotGemini CLI, OpenCode, and OpenRouter

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadGemini 3.8 FlashGemini 3.5 Flash-Lite
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$0.71$0.34
Large one-off review, 150K input with no cache hits, 10K output$0.15$0.07
Output-heavy generation, 30K input, 80K output$0.32$0.21
A month of sessions, 110 sessions: 5 a day, 22 working days$78.38$36.85
Where the session’s cost goes
Cache writes$0.30$0.12
Cache reads$0.15$0.06
Uncached input$0.08$0.03
Output$0.19$0.13
caching saves on the session with Gemini 3.8 Flash (66%)
$1.35
caching saves on the session with Gemini 3.5 Flash-Lite (61%)
$0.54

What Google recommends each model for

Google positions Gemini 3.8 Flash as its most intelligent Flash, aimed at long-horizon software engineering, autonomous agents, and enterprise workflows. It claims complex multi-file refactoring, deterministic tool execution, and multi-step planning with fewer failed loops.

Gemini 3.5 Flash-Lite is Google's low-cost, low-latency tier for high-throughput execution. Google names subagent tasks and document parsing as its uses, cites about 350 output tokens per second, and includes computer use as a built-in tool. Its default thinking level is minimal, against medium on 3.8 Flash.

Google recommends the two together for new projects. Read alongside the positioning, that suggests a main agent on 3.8 Flash handing narrow, repeated steps to Flash-Lite, rather than a choice of one over the other.

Why output-heavy work narrows the price gap

Input and caching sit 2.5x apart: $0.75 against $0.30 per million input tokens, and $0.075 against $0.03 for cache hits. Output is closer, $3.75 against $2.50, or 1.5x. Google bills no separate cache write on either, so written tokens cost ordinary input.

So the gap depends on the work. Input-heavy work shows most of the gap: $0.15 against $0.07 for the uncached review and $0.71 against $0.34 for the session, roughly 2.1x each. The output-heavy generation costs $0.32 against $0.21, only 1.5x. Output is 38% of Flash-Lite's session cost but 26% of 3.8 Flash's.

Over 110 sessions a month, 3.8 Flash comes to $78.38 and Flash-Lite to $36.85, a difference of $41.53. Caching saves $1.35 on 3.8 Flash, or 66%, and $0.54 on Flash-Lite, or 61%.

Introductory pricing, Gemini CLI, and where each runs

Gemini 3.8 Flash's rates are introductory through December 31, 2026. Starting January 1, 2027, its rates become $1.50 for input, $0.15 for cache hits, and $7.50 for output. Gemini 3.5 Flash-Lite's rates carry no introductory note, so the figures above describe the gap for the rest of 2026.

In Gemini CLI, 3.8 Flash is the Flash half of the default auto model for Gemini API key and Vertex AI users, and the CLI sends it a high thinking level. The flash-lite alias resolves to 3.5 Flash-Lite. Both run in OpenCode and on OpenRouter, while Cursor and GitHub Copilot list 3.8 Flash only.

On limits they tie, at a 1.05M context window and 65.5K of output. 3.8 Flash also has free-tier coverage on the Gemini API for input, output, and caching. EveryToken reads your local Gemini CLI history and prices each request at Google's rates, so you can see how much of a session ran on each model.

Prompt caching

How Google bills cached tokens

Google

Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.

Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.

On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.

Source: Google: Context caching

Your own numbers

See what Gemini 3.8 Flash and Gemini 3.5 Flash-Lite really cost you.

everyaitoken reads your Gemini CLI, Cursor, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is Gemini 3.5 Flash-Lite cheaper than Gemini 3.8 Flash?

Yes. It charges $0.30 input and $2.50 output per million tokens against $0.75 and $3.75. The example session costs $0.34 against $0.71.

Which model does Gemini CLI use for flash-lite?

The flash-lite alias in Gemini CLI resolves to Gemini 3.5 Flash-Lite. The default auto model uses Gemini 3.8 Flash as its Flash half for Gemini API key and Vertex AI users.

Will Gemini 3.8 Flash get more expensive?

Yes, once its introductory rates end after December 31, 2026. The 2027 price list has it at $1.50 input, $0.15 cached, and $7.50 output per million tokens.

Is Gemini 3.5 Flash-Lite in Cursor?

No. It runs in Gemini CLI, OpenCode, and OpenRouter, but not in Cursor or GitHub Copilot.