Model comparison
Gemini 3.1 Pro Preview vs Gemini 2.5 Pro: preview or stable?
Gemini 3.1 Pro Preview costs 60% more per input token than Gemini 2.5 Pro and is still a preview, while 2.5 Pro now limits new access. How to choose.
· Prices as of September 28, 2026
Gemini 3.1 Pro Preview
Google · Released February 19, 2026 · Preview
Google's only current Pro model, still in preview, positioned for deep reasoning and agentic coding.
Gemini 3.1 Pro Preview facts and comparisonsGemini 2.5 Pro
Google · Released June 17, 2025 · Previous generation
The previous-generation Pro, still stable, and still Gemini CLI's Pro fallback for accounts without preview access.
Gemini 2.5 Pro facts and comparisons
The short answer
Gemini 3.1 Pro Preview is Google's current Pro model and costs more than Gemini 2.5 Pro: the example agentic coding session is $2.00 against $1.38. Gemini 2.5 Pro is the stable previous generation, but since September 18, 2026, Google limits it to accounts that used it before. Between the two, new accounts can only choose the Pro Preview, while accounts already on 2.5 Pro can keep a stable model at lower rates.
Choose Gemini 3.1 Pro Preview if
- Your account never used Gemini 2.5 Pro, so Google's access limit leaves 3.1 Pro Preview as your Pro option.
- You want Gemini CLI's default auto model, where 3.1 Pro Preview is the Pro half.
- You use Cursor, which lists 3.1 Pro Preview but not Gemini 2.5 Pro.
- Your agents lean on custom tools alongside bash, which the customtools endpoint prioritizes.
Choose Gemini 2.5 Pro if
- Your account already uses Gemini 2.5 Pro and you want a stable, non-preview model.
- You want the lower rates: $1.25 input and $10 output per million, against $2 and $12.
- You prefer setting thinking as a token budget rather than choosing a thinking level.
Side by side
Specs and prices
| Fact | Gemini 3.1 Pro Preview | Gemini 2.5 Pro |
|---|---|---|
| Maker | ||
| API model id | gemini-3.1-pro-preview | gemini-2.5-pro |
| Released | February 19, 2026 | June 17, 2025 |
| Status | Preview | Previous generation |
| Context window | 1.05M tokens | 1.05M tokens |
| Max output | 65.5K tokens | 65.5K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $2 | $1.25 |
| Cache hit, per 1M | $0.20 | $0.125 |
| Cache write, per 1M | $2 (same as input) | $1.25 (same as input) |
| Output, per 1M tokens | $12 | $10 |
| Runs in | Cursor, Gemini CLI, OpenCode, and OpenRouter | Gemini CLI, OpenCode, and OpenRouter |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Gemini 3.1 Pro Preview: Prompts over 200K input tokens cost $4 input, $0.40 cached, and $18 output per million tokens. Gemini 2.5 Pro: Prompts over 200K input tokens cost $2.50 input, $0.25 cached, and $15 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Gemini 3.1 Pro Preview | Gemini 2.5 Pro |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $2.00 | $1.38 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.42 | $0.29 |
| Output-heavy generation, 30K input, 80K output | $1.02 | $0.84 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $220.00 | $151.25 |
| Where the session’s cost goes | ||
| Cache writes | $0.80 | $0.50 |
| Cache reads | $0.40 | $0.25 |
| Uncached input | $0.20 | $0.13 |
| Output | $0.60 | $0.50 |
- caching saves on the session with Gemini 3.1 Pro Preview (64%)
- $3.60
- caching saves on the session with Gemini 2.5 Pro (62%)
- $2.25
How much more does Gemini 3.1 Pro Preview cost?
Gemini 3.1 Pro Preview charges $2 per million input tokens, $0.20 per million cached, and $12 per million output. Gemini 2.5 Pro charges $1.25, $0.125, and $10. Input costs 60% more on the Pro Preview, while output costs only 20% more.
Because the input gap is the larger one, input-heavy work shows the bigger difference. The example session and the large one-off review both cost 45% more on the Pro Preview: $2.00 against $1.38, and $0.42 against $0.29. The output-heavy generation costs 21% more, $1.02 against $0.84. Over 110 sessions a month, the estimate is $220.00 against $151.25 at API rates.
Neither model has a separate cache-write price. Google bills written tokens as ordinary input, and a hit costs 10% of input on both, so caching saves a similar share on each, 64% on the Pro Preview and 62% on 2.5 Pro. The cache doesn't change the ratio between them much.
What happens above 200K input tokens
Both models raise their rates for prompts over 200K input tokens. The Pro Preview goes to $4 input, $0.40 cached, and $18 output per million. Gemini 2.5 Pro goes to $2.50 input, $0.25 cached, and $15 output. Both context windows are 1.05M tokens with up to 65.5K of output, so an agent that fills more than about a fifth of the window pays the higher tier.
The example session keeps every request under 200K, so the tables show the standard tier. If your agent sends large prompts regularly, compare the upper tier instead: input still costs 60% more on the Pro Preview there, and output is $18 against $15.
Preview status, access limits, and Gemini CLI
Gemini 3.1 Pro Preview is Google's current Pro model, and it is still a preview. Its announced successor, Gemini 3.5 Pro, has not been released yet. Google names software engineering and agentic workflows that need precise tool use and reliable multi-step execution as the Pro Preview's strengths, and offers a separate customtools endpoint that is better at prioritizing custom tools alongside bash.
Gemini 2.5 Pro is the stable previous generation. Google calls it "A Pro model which excels at coding and complex reasoning tasks," and it is not deprecated. But since September 18, 2026, Google limits access to accounts that used it before, so it is closed to new users.
In Gemini CLI, 3.1 Pro Preview is the Pro half of the default auto model, and Gemini API key users get the customtools endpoint at the same price. Gemini 2.5 Pro is the Pro fallback for accounts without preview access. Thinking works differently: the Pro Preview's thinking can't be turned off and defaults to high, while 2.5 Pro uses thinking budgets, and Gemini CLI sends 8,192 tokens.
Outside Gemini CLI, the Pro Preview is in Cursor, OpenRouter, and OpenCode, and GitHub Copilot retired it on September 1, 2026. Gemini 2.5 Pro is in OpenRouter and OpenCode but not Cursor.
Prompt caching
How Google bills cached tokens
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what Gemini 3.1 Pro Preview and Gemini 2.5 Pro really cost you.
everyaitoken reads your Gemini CLI, Cursor, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is Gemini 3.1 Pro Preview more expensive than Gemini 2.5 Pro?
Yes. Input costs $2 against $1.25 per million and output $12 against $10. The example agentic session is $2.00 against $1.38.
Can new users still access Gemini 2.5 Pro?
No. Since September 18, 2026, Google limits it to accounts that used it before. It is not deprecated.
Is Gemini 3.1 Pro Preview a stable model?
No, it is still a preview. Its successor, Gemini 3.5 Pro, is announced but has no release yet.
How can I see what a move between the two changes?
EveryToken reads your local Gemini CLI, Cursor, and OpenCode history on your Mac and prices every request at Google's API rates by model, with what caching saved on each. Requests made through OpenRouter are priced from OpenRouter's catalog.
Sources
- Google: Gemini API pricing
- Google docs: Gemini 3.1 Pro Preview
- Google: Gemini models
- Google: Gemini 3.1 Pro
- Google DeepMind: Gemini
- Gemini CLI source: model configuration
- Cursor docs: Gemini 3.1 Pro
- OpenRouter: Gemini 3.1 Pro Preview
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Google docs: Gemini 2.5 Pro
- Google: Gemini API changelog
- OpenRouter: Gemini 2.5 Pro
- Google: Context caching