Skip to content

Model comparison

Replacing Gemini 3.1 Flash-Lite with Gemini 3.5 Flash-Lite

Gemini 3.1 Flash-Lite shuts down on May 7, 2027, and Gemini 3.5 Flash-Lite replaces it at higher rates. What the move costs, mostly on output.

· Prices as of September 28, 2026

  • Gemini 3.5 Flash-Lite

    Google · Released July 21, 2026

    Google's current low-cost, low-latency tier for high-volume and subagent work, recommended with Gemini 3.8 Flash for new projects.

    Gemini 3.5 Flash-Lite facts and comparisons
  • Gemini 3.1 Flash-Lite

    Google · Released May 7, 2026 · Previous generation

    The older budget tier for lightweight, high-frequency tasks such as routing and extraction, superseded by Gemini 3.5 Flash-Lite.

    Gemini 3.1 Flash-Lite facts and comparisons

The short answer

Gemini 3.5 Flash-Lite is Google's named replacement for Gemini 3.1 Flash-Lite, which shuts down on May 7, 2027. The newer model costs 20% more for input and 67% more for output, so the example agentic coding session rises from $0.25 to $0.34. Google positions 3.5 Flash-Lite for subagent tasks and high-throughput execution, while 3.1 Flash-Lite was built for lightweight, high-frequency jobs such as routing and extraction.

Choose Gemini 3.5 Flash-Lite if

  • You want the model Google names as the replacement, ahead of the May 7, 2027 shutdown of 3.1 Flash-Lite.
  • You run subagent tasks or document parsing, which Google names for 3.5 Flash-Lite.
  • You use Gemini CLI's flash-lite alias, which now resolves to 3.5 Flash-Lite.
  • You want computer use as a built-in tool for agentic tasks.

Choose Gemini 3.1 Flash-Lite if

  • You run routing and classification at high volume and want the lower rate until the shutdown: $0.25 input and $1.50 output per million tokens.
  • Your jobs are output-heavy, where the example generation costs $0.13 against $0.21.
  • You need time to validate the replacement on real traffic before May 7, 2027.

Side by side

Specs and prices

FactGemini 3.5 Flash-LiteGemini 3.1 Flash-Lite
MakerGoogleGoogle
API model idgemini-3.5-flash-litegemini-3.1-flash-lite
ReleasedJuly 21, 2026May 7, 2026
StatusCurrentPrevious generation
Context window1.05M tokens1.05M tokens
Max output65.5K tokens65.5K tokens
Open weightsNoNo
Input, per 1M tokens$0.30$0.25
Cache hit, per 1M$0.03$0.025
Cache write, per 1M$0.30 (same as input)$0.25 (same as input)
Output, per 1M tokens$2.50$1.50
Runs inGemini CLI, OpenCode, and OpenRouterGemini CLI, OpenCode, and OpenRouter

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadGemini 3.5 Flash-LiteGemini 3.1 Flash-Lite
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$0.34$0.25
Large one-off review, 150K input with no cache hits, 10K output$0.07$0.05
Output-heavy generation, 30K input, 80K output$0.21$0.13
A month of sessions, 110 sessions: 5 a day, 22 working days$36.85$27.50
Where the session’s cost goes
Cache writes$0.12$0.10
Cache reads$0.06$0.05
Uncached input$0.03$0.03
Output$0.13$0.08
caching saves on the session with Gemini 3.5 Flash-Lite (61%)
$0.54
caching saves on the session with Gemini 3.1 Flash-Lite (64%)
$0.45

When does Gemini 3.1 Flash-Lite shut down?

Google lists Gemini 3.1 Flash-Lite for shutdown on May 7, 2027, with Gemini 3.5 Flash-Lite as its replacement. Anything that calls the older model needs to move before then.

Gemini 3.1 Flash-Lite launched on May 7, 2026, as a low-latency, cost-effective multimodal model for high-frequency, lightweight tasks. Google cites model routing and classification as uses, noting that Gemini CLI uses Flash-Lite to route requests to Flash or Pro. Gemini 3.5 Flash-Lite, released on July 21, 2026, is Google's current low-cost tier, recommended with Gemini 3.8 Flash for new projects.

What the replacement costs, mostly on output

Input rises modestly, from $0.25 to $0.30 per million tokens, 20% more, with cache hits moving from $0.025 to $0.03. Output rises more steeply, from $1.50 to $2.50 per million tokens, 67% more. Neither has a separate cache-write price, so Google charges written tokens as ordinary input.

The workloads reflect that split. The uncached review costs $0.05 on 3.1 Flash-Lite and $0.07 on 3.5 Flash-Lite. The session costs $0.25 against $0.34, and the output-heavy generation $0.13 against $0.21. Output is 38% of 3.5 Flash-Lite's session and 31% of 3.1 Flash-Lite's.

Over 110 sessions a month, that is $27.50 on the older model and $36.85 on the newer one, $9.35 more. The difference is small per developer and grows with volume, which is how Flash-Lite models tend to be used. Routing and classification calls usually return short outputs, so they would feel the 20% input increase more than the 67% output one.

Thinking levels, Gemini CLI, and where each runs

Gemini 3.5 Flash-Lite's default thinking level is minimal. Google cites about 350 output tokens per second for it and lists computer use as a built-in tool for agentic tasks. Thinking level changes how many tokens a request writes, so raising it on 3.5 Flash-Lite brings its higher output rate into play.

In Gemini CLI, the flash-lite alias resolves to 3.5 Flash-Lite. Both models run in Gemini CLI, OpenCode, and OpenRouter, and neither is in Cursor or GitHub Copilot. Their limits are the same as well: 1.05M tokens of context and 65.5K of output.

Before May 7, 2027, run the replacement on real traffic and compare output lengths. EveryToken reads your local Gemini CLI and OpenCode history and prices each request at Google's rates, split by model, and prices requests made through OpenRouter from OpenRouter's catalog.

Prompt caching

How Google bills cached tokens

Google

Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.

Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.

On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.

Source: Google: Context caching

Your own numbers

See what Gemini 3.5 Flash-Lite and Gemini 3.1 Flash-Lite really cost you.

everyaitoken reads your Gemini CLI, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

What replaces Gemini 3.1 Flash-Lite?

Google names Gemini 3.5 Flash-Lite as the replacement. Gemini 3.1 Flash-Lite shuts down on May 7, 2027.

Is Gemini 3.5 Flash-Lite more expensive than Gemini 3.1 Flash-Lite?

Yes. It charges $0.30 input and $2.50 output per million tokens against $0.25 and $1.50. A month of 110 example sessions costs $36.85 against $27.50.

Which Flash-Lite model does Gemini CLI use?

Gemini CLI's flash-lite alias resolves to Gemini 3.5 Flash-Lite. Both Flash-Lite models are available in Gemini CLI.

Do both Flash-Lite models support caching?

Yes. On both, a cache hit costs 10% of the input price, $0.03 on 3.5 Flash-Lite and $0.025 on 3.1 Flash-Lite. Explicit caching adds a storage charge of $0.50 to $1 per million tokens per hour on Flash models.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash-Lite

    Google recommends Gemini 3.8 Flash and Gemini 3.5 Flash-Lite together for new projects. Where each fits, the 2.1x session gap, and what changes in 2027.

  • Claude Haiku 4.5 vs Gemini 3.5 Flash-Lite

    Claude Haiku 4.5 and Gemini 3.5 Flash-Lite both target sub-agent work. Haiku 4.5 costs 3.5x as much per cached session, but only 2x as much per output token.

  • GPT-5.6 Luna vs Gemini 3.5 Flash-Lite

    Gemini 3.5 Flash-Lite costs 1.5x as much as GPT-5.6 Luna on a cached coding session, $0.34 against $0.22, and about twice as much on output-heavy work.

  • GPT-6 Luna vs Gemini 3.5 Flash-Lite

    GPT-6 Luna lists a third of Gemini 3.5 Flash-Lite's input rate and a fifth of its output rate. A month of cached coding sessions: $11.55 against $36.85.

  • Gemini 2.5 Pro vs Gemini 2.5 Flash

    Gemini 2.5 Pro costs about 4x Gemini 2.5 Flash, and since September 18, 2026, only earlier users can reach either. What each costs and where to go next.

  • Gemini 3.1 Pro Preview vs Gemini 2.5 Pro

    Gemini 3.1 Pro Preview costs 60% more per input token than Gemini 2.5 Pro and is still a preview, while 2.5 Pro now limits new access. How to choose.