Gemini 3.5 Flash-Lite: price, context window, and caching
Google's current low-cost, low-latency tier for high-volume and subagent work, recommended with Gemini 3.8 Flash for new projects.
Released July 21, 2026 · Prices as of September 28, 2026
In Google’s words
“Our fastest, most cost-effective 3.5 model for high-throughput execution.”
Facts
Specs and prices
| Fact | Gemini 3.5 Flash-Lite |
|---|---|
| Maker | |
| API model id | gemini-3.5-flash-lite |
| Released | July 21, 2026 |
| Status | Current |
| Context window | 1.05M tokens |
| Max output | 65.5K tokens |
| Open weights | No |
| Input, per 1M tokens | $0.30 |
| Cache hit, per 1M | $0.03 |
| Cache write, per 1M | $0.30 (same as input) |
| Output, per 1M tokens | $2.50 |
| Runs in | Gemini CLI, OpenCode, and OpenRouter |
Standard API rates in US dollars, as published by the maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included.
Good to know
- Gemini CLI uses it for the flash-lite alias.
- The default thinking level is minimal.
Cost
What typical work costs
Example token counts at Gemini 3.5 Flash-Lite’s published rates. On the agentic session, caching saves $0.54 against billing every token as ordinary input.
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.34 |
|---|---|
| Large one-off review, 150K input with no cache hits, 10K output | $0.07 |
| Output-heavy generation, 30K input, 80K output | $0.21 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $36.85 |
Prompt caching
How Google bills cached tokens
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Gemini 3.5 Flash-Lite compared
Claude Haiku 4.5 vs Gemini 3.5 Flash-Lite
Claude Haiku 4.5 and Gemini 3.5 Flash-Lite both target sub-agent work. Haiku 4.5 costs 3.5x as much per cached session, but only 2x as much per output token.
Gemini 3.8 Flash vs Gemini 3.5 Flash-Lite
Google recommends Gemini 3.8 Flash and Gemini 3.5 Flash-Lite together for new projects. Where each fits, the 2.1x session gap, and what changes in 2027.
GPT-5.6 Luna vs Gemini 3.5 Flash-Lite
Gemini 3.5 Flash-Lite costs 1.5x as much as GPT-5.6 Luna on a cached coding session, $0.34 against $0.22, and about twice as much on output-heavy work.
GPT-6 Luna vs Gemini 3.5 Flash-Lite
GPT-6 Luna lists a third of Gemini 3.5 Flash-Lite's input rate and a fifth of its output rate. A month of cached coding sessions: $11.55 against $36.85.
Gemini 3.5 Flash-Lite vs Gemini 3.1 Flash-Lite
Gemini 3.1 Flash-Lite shuts down on May 7, 2027, and Gemini 3.5 Flash-Lite replaces it at higher rates. What the move costs, mostly on output.
Your own numbers
See what Gemini 3.5 Flash-Lite really costs you.
everyaitoken reads your Gemini CLI, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.