Gemini 2.5 Flash: price, context window, and caching
The previous-generation price-performance Flash with controllable thinking budgets.
Released June 17, 2025 · Prices as of September 28, 2026
In Google’s words
“Our best price-performance model for low-latency, high-volume tasks that require reasoning.”
What Google says it’s good at
- Low-latency, high-volume tasks that need thinking, and agentic use Source
Facts
Specs and prices
| Fact | Gemini 2.5 Flash |
|---|---|
| Maker | |
| API model id | gemini-2.5-flash |
| Released | June 17, 2025 |
| Status | Previous generation |
| Context window | 1.05M tokens |
| Max output | 65.5K tokens |
| Open weights | No |
| Input, per 1M tokens | $0.30 |
| Cache hit, per 1M | $0.03 |
| Cache write, per 1M | $0.30 (same as input) |
| Output, per 1M tokens | $2.50 |
| Runs in | Gemini CLI, OpenCode, and OpenRouter |
Standard API rates in US dollars, as published by the maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included.
Good to know
- Since September 18, 2026, Google limits access to accounts that used it before. It is not deprecated.
Cost
What typical work costs
Example token counts at Gemini 2.5 Flash’s published rates. On the agentic session, caching saves $0.54 against billing every token as ordinary input.
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.34 |
|---|---|
| Large one-off review, 150K input with no cache hits, 10K output | $0.07 |
| Output-heavy generation, 30K input, 80K output | $0.21 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $36.85 |
Prompt caching
How Google bills cached tokens
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Gemini 2.5 Flash compared
Gemini 2.5 Pro vs Gemini 2.5 Flash
Gemini 2.5 Pro costs about 4x Gemini 2.5 Flash, and since September 18, 2026, only earlier users can reach either. What each costs and where to go next.
Your own numbers
See what Gemini 2.5 Flash really costs you.
everyaitoken reads your Gemini CLI, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.