Skip to content

Google

Gemini 3.5 Flash: price, context window, and caching

Launched in May 2026 as Google's agent and coding Flash, now called its legacy Flash for routine, high-throughput work.

Released May 19, 2026 · Prices as of September 28, 2026

In Google’s words

“Our legacy Flash model, providing baseline speed and foundational performance for routine, high-throughput workloads.”

Google: Gemini models

What Google says it’s good at

  • Output four times faster than other frontier models, by Google's measure Source
  • Sub-agent deployment, multi-step workflows, and long-horizon tasks at scale Source

Facts

Specs and prices

FactGemini 3.5 Flash
MakerGoogle
API model idgemini-3.5-flash
ReleasedMay 19, 2026
StatusPrevious generation
Context window1.05M tokens
Max output65.5K tokens
Open weightsNo
Input, per 1M tokens$1.50
Cache hit, per 1M$0.15
Cache write, per 1M$1.50 (same as input)
Output, per 1M tokens$9
Runs inGemini CLI, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by the maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included.

Good to know

  • While Gemini 3.6 to 3.8 Flash are on introductory rates, it costs more per token than they do.
  • Gemini CLI uses it as the Flash model for Sign in with Google accounts that don't yet get the latest Flash.

Cost

What typical work costs

Example token counts at Gemini 3.5 Flash’s published rates. On the agentic session, caching saves $2.70 against billing every token as ordinary input.

Example workload costs for Gemini 3.5 Flash
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$1.50
Large one-off review, 150K input with no cache hits, 10K output$0.32
Output-heavy generation, 30K input, 80K output$0.77
A month of sessions, 110 sessions: 5 a day, 22 working days$165.00

Prompt caching

How Google bills cached tokens

Google

Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.

Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.

On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.

Source: Google: Context caching

Gemini 3.5 Flash compared

  • Claude Sonnet 5 vs Gemini 3.5 Flash

    Gemini 3.5 Flash, now Google's legacy Flash, costs $1.50 against $2.40 on Claude Sonnet 5 per cached coding session. Most of the gap is cache writes.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash

    Gemini 3.8 Flash costs about half of Gemini 3.5 Flash on introductory rates. What changes on January 1, 2027, and which one Gemini CLI picks for you.

Your own numbers

See what Gemini 3.5 Flash really costs you.

everyaitoken reads your Gemini CLI, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math