Model comparison
Gemini 2.5 Pro vs Gemini 2.5 Flash after the access limit
Gemini 2.5 Pro costs about 4x Gemini 2.5 Flash, and since September 18, 2026, only earlier users can reach either. What each costs and where to go next.
· Prices as of September 28, 2026
Gemini 2.5 Pro
Google · Released June 17, 2025 · Previous generation
The previous-generation Pro, still stable, and still Gemini CLI's Pro fallback for accounts without preview access.
Gemini 2.5 Pro facts and comparisonsGemini 2.5 Flash
Google · Released June 17, 2025 · Previous generation
The previous-generation price-performance Flash with controllable thinking budgets.
Gemini 2.5 Flash facts and comparisons
The short answer
Gemini 2.5 Pro costs about 4x Gemini 2.5 Flash on every rate, so the example agentic coding session comes to $1.38 against $0.34. Since September 18, 2026, Google limits access to both to accounts that used them before, though neither is deprecated. If you still have access, Pro is the one Google pitches for complex reasoning in code and Gemini CLI's Pro fallback, and Flash is its price-performance option for low-latency, high-volume tasks.
Choose Gemini 2.5 Pro if
- You reason over complex problems in code, math, and STEM, which Google names as Gemini 2.5 Pro's strength.
- You analyze large codebases and documents with long context, as Google describes for 2.5 Pro.
- Your Gemini CLI account lacks preview access, and 2.5 Pro is the Pro fallback there.
Choose Gemini 2.5 Flash if
- You run low-latency, high-volume tasks that need thinking, the work Google names for Gemini 2.5 Flash.
- You want the lower rates: $0.30 input and $2.50 output per million tokens.
- You send long prompts, and the rates this page uses list no higher tier for 2.5 Flash above 200K input tokens.
Side by side
Specs and prices
| Fact | Gemini 2.5 Pro | Gemini 2.5 Flash |
|---|---|---|
| Maker | ||
| API model id | gemini-2.5-pro | gemini-2.5-flash |
| Released | June 17, 2025 | June 17, 2025 |
| Status | Previous generation | Previous generation |
| Context window | 1.05M tokens | 1.05M tokens |
| Max output | 65.5K tokens | 65.5K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $1.25 | $0.30 |
| Cache hit, per 1M | $0.125 | $0.03 |
| Cache write, per 1M | $1.25 (same as input) | $0.30 (same as input) |
| Output, per 1M tokens | $10 | $2.50 |
| Runs in | Gemini CLI, OpenCode, and OpenRouter | Gemini CLI, OpenCode, and OpenRouter |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Gemini 2.5 Pro: Prompts over 200K input tokens cost $2.50 input, $0.25 cached, and $15 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Gemini 2.5 Pro | Gemini 2.5 Flash |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $1.38 | $0.34 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.29 | $0.07 |
| Output-heavy generation, 30K input, 80K output | $0.84 | $0.21 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $151.25 | $36.85 |
| Where the session’s cost goes | ||
| Cache writes | $0.50 | $0.12 |
| Cache reads | $0.25 | $0.06 |
| Uncached input | $0.13 | $0.03 |
| Output | $0.50 | $0.13 |
- caching saves on the session with Gemini 2.5 Pro (62%)
- $2.25
- caching saves on the session with Gemini 2.5 Flash (61%)
- $0.54
Can you still use Gemini 2.5 Pro and Gemini 2.5 Flash?
Since September 18, 2026, Google limits access to Gemini 2.5 Pro and Gemini 2.5 Flash to accounts that used them before. Neither is deprecated. If your account never called them, this comparison is mostly useful for understanding older usage and what it cost.
Both were released on June 17, 2025, and Google's lineup has moved on since. Gemini 3.8 Flash is its newest Flash, and Gemini 3.1 Pro Preview is its current Pro, still in preview. Gemini 2.5 Pro remains Gemini CLI's Pro fallback for accounts without preview access.
How much more does Gemini 2.5 Pro cost?
For every million tokens, Gemini 2.5 Pro bills $1.25 of input, $0.125 of cache hits, and $10 of output, where Gemini 2.5 Flash bills $0.30, $0.03, and $2.50. Input and cache hits are 4.2x apart, and output 4x.
The workloads land close to 4x throughout. The session costs $1.38 against $0.34, the uncached review $0.29 against $0.07, and the output-heavy generation $0.84 against $0.21. Over 110 sessions a month, Pro comes to $151.25 and Flash to $36.85, a difference of $114.40.
Output and cache writes each take 36% of Pro's session, since Google bills written tokens as ordinary input. Caching saves $2.25 on Pro, or 62%, and $0.54 on Flash, or 61%.
Long prompts, explicit caching, and thinking budgets
Gemini 2.5 Pro has a long-context tier: prompts over 200K input tokens cost $2.50 input, $0.25 cached, and $15 output per million tokens. The rates this page uses list no such tier for Gemini 2.5 Flash, so long prompts widen the gap. None of the example session's requests reach that size.
Explicit caching costs more on Pro as well. Google adds a storage charge while an explicit cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models, and $0.50 to $1 on Flash models. Implicit caching is on by default for Gemini 2.5 and newer and applies the discount automatically, though a hit is not assured.
Both models use thinking budgets rather than thinking levels, and Gemini CLI sends Gemini 2.5 Pro a budget of 8,192 tokens. Thinking changes output length, and output is 36% of Pro's session cost, so the budget is a real cost lever. EveryToken reads your local Gemini CLI history and prices each request at Google's rates, so you can see what your remaining 2.5 usage costs.
Prompt caching
How Google bills cached tokens
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what Gemini 2.5 Pro and Gemini 2.5 Flash really cost you.
everyaitoken reads your Gemini CLI, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is Gemini 2.5 Pro deprecated?
No. Since September 18, 2026, Google limits access to accounts that used it before, but it is not deprecated. The same applies to Gemini 2.5 Flash.
How much cheaper is Gemini 2.5 Flash than Gemini 2.5 Pro?
It costs $0.30 input and $2.50 output per million tokens, against $1.25 and $10 for Pro. The example session costs $0.34 on Flash and $1.38 on Pro, 75% less.
Does Gemini CLI still use Gemini 2.5 Pro?
Yes, as the Pro fallback for accounts without preview access. Gemini CLI sends it a thinking budget of 8,192 tokens.
Does Gemini 2.5 Pro charge more for long prompts?
Yes. Prompts over 200K input tokens cost $2.50 input, $0.25 cached, and $15 output per million tokens.