Gemini 3.1 Pro Preview: price, context window, and caching
Google's only current Pro model, still in preview, positioned for deep reasoning and agentic coding.
Released February 19, 2026 · Prices as of September 28, 2026
In Google’s words
“Advanced intelligence, complex problem-solving skills, and powerful agentic and vibe coding capabilities.”
Facts
Specs and prices
| Fact | Gemini 3.1 Pro Preview |
|---|---|
| Maker | |
| API model id | gemini-3.1-pro-preview |
| Released | February 19, 2026 |
| Status | Preview |
| Context window | 1.05M tokens |
| Max output | 65.5K tokens |
| Open weights | No |
| Input, per 1M tokens | $2 |
| Cache hit, per 1M | $0.20 |
| Cache write, per 1M | $2 (same as input) |
| Output, per 1M tokens | $12 |
| Runs in | Cursor, Gemini CLI, OpenCode, and OpenRouter |
Standard API rates in US dollars, as published by the maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Gemini 3.1 Pro Preview: Prompts over 200K input tokens cost $4 input, $0.40 cached, and $18 output per million tokens.
Good to know
- The Pro half of Gemini CLI's default auto model. Gemini API key users get the gemini-3.1-pro-preview-customtools endpoint, at the same price.
- Google has announced Gemini 3.5 Pro, which is not yet released.
- Thinking cannot be turned off, and the default level is high.
- GitHub Copilot retired it on September 1, 2026.
Cost
What typical work costs
Example token counts at Gemini 3.1 Pro Preview’s published rates. On the agentic session, caching saves $3.60 against billing every token as ordinary input.
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $2.00 |
|---|---|
| Large one-off review, 150K input with no cache hits, 10K output | $0.42 |
| Output-heavy generation, 30K input, 80K output | $1.02 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $220.00 |
Prompt caching
How Google bills cached tokens
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Gemini 3.1 Pro Preview compared
Claude Fable 5.1 vs Gemini 3.1 Pro Preview
A cached coding session costs 5.3x more on Claude Fable 5.1 than on Gemini 3.1 Pro Preview, mostly from cache writes. Output limits and preview status compared.
Claude Opus 4.8 vs Gemini 3.1 Pro Preview
Claude Opus 4.8 costs 3x Gemini 3.1 Pro Preview on a cached coding session, $6.00 against $2.00. How long prompts, fast mode, and preview status shift that.
Claude Opus 5.5 vs Gemini 3.1 Pro Preview
Gemini 3.1 Pro Preview costs less than half of Claude Opus 5.5 on a cached coding session, though both charge $0.20 per cache hit. Where the gap comes from.
Claude Sonnet 5 vs Gemini 3.1 Pro Preview
Claude Sonnet 5 and Gemini 3.1 Pro Preview both charge $2 input, but Gemini's rates rise past 200K tokens and it is still a preview. Session: $2.40 vs $2.00.
Gemini 3.1 Pro Preview vs Gemini 2.5 Pro
Gemini 3.1 Pro Preview costs 60% more per input token than Gemini 2.5 Pro and is still a preview, while 2.5 Pro now limits new access. How to choose.
Gemini 3.8 Flash vs Gemini 3.1 Pro Preview
Gemini CLI's auto model uses both Gemini 3.8 Flash and Gemini 3.1 Pro Preview. What each costs, why Flash rates double in 2027, and the Pro's preview status.
GLM-5.3 vs Gemini 3.1 Pro Preview
Neither Z.ai nor Google charges extra to write the cache, so GLM-5.3 and Gemini 3.1 Pro Preview split on output: $4.40 against $12 per million tokens.
GPT-5.3-Codex vs Gemini 3.1 Pro Preview
GPT-5.3-Codex and Gemini 3.1 Pro Preview cost within 4% of each other on a cached coding session. Context size, output limits, and access set them apart.
GPT-5.5 vs Gemini 3.1 Pro Preview
GPT-5.5 costs 2.5x as much as Gemini 3.1 Pro Preview on every rate and every workload. Where they differ instead: output limits, long prompts, and access.
GPT-5.6 Sol vs Gemini 3.1 Pro Preview
GPT-5.6 Sol is on promotional rates and Gemini 3.1 Pro Preview is still in preview. On a cached coding session, Gemini costs $2.00 against $4.20.
GPT-6 Astra vs Gemini 3.1 Pro Preview
GPT-6 Astra costs 5.3x as much as Gemini 3.1 Pro Preview on a cached coding session. How OpenAI's write premium, long-context tiers, and output caps compare.
GPT-6 Sol vs Gemini 3.1 Pro Preview
GPT-6 Sol and Gemini 3.1 Pro Preview both charge $2 input and $0.20 per cache hit. Which one costs less flips by workload, so limits and status decide.
Grok 4.7 vs Gemini 3.1 Pro Preview
Grok 4.7 and Gemini 3.1 Pro Preview both charge $2 input and raise rates at 200K tokens. Gemini costs less on a cached session, Grok on output-heavy work.
Your own numbers
See what Gemini 3.1 Pro Preview really costs you.
everyaitoken reads your Gemini CLI, Cursor, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
Sources
- Google: Gemini API pricing
- Google docs: Gemini 3.1 Pro Preview
- Google: Gemini models
- Google: Gemini 3.1 Pro
- Google DeepMind: Gemini
- Gemini CLI source: model configuration
- Cursor docs: Gemini 3.1 Pro
- OpenRouter: Gemini 3.1 Pro Preview
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Google: Context caching