Model comparison
Claude Opus 5.5 vs Gemini 3.1 Pro Preview: cost and limits
Gemini 3.1 Pro Preview costs less than half of Claude Opus 5.5 on a cached coding session, though both charge $0.20 per cache hit. Where the gap comes from.
· Prices as of September 28, 2026
Claude Opus 5.5
Anthropic · Released September 22, 2026
Anthropic's recommended starting model for most work, built for long-running agentic coding. Anthropic says it matches Claude Fable 5.1 on most work at a much lower price.
Claude Opus 5.5 facts and comparisonsGemini 3.1 Pro Preview
Google · Released February 19, 2026 · Preview
Google's only current Pro model, still in preview, positioned for deep reasoning and agentic coding.
Gemini 3.1 Pro Preview facts and comparisons
The short answer
Gemini 3.1 Pro Preview is cheaper: the example agentic coding session costs $2.00 on it and $4.40 on Claude Opus 5.5, a 2.2x gap, even though both charge $0.20 per million for a cache hit. Opus 5.5 is Claude Code's default and Anthropic's recommended starting model, with 128K of output and no long-context surcharge. Gemini 3.1 Pro Preview fits Gemini CLI users who want Google's Pro model at a lower price and can work with a preview.
Choose Claude Opus 5.5 if
- You work in Claude Code, where Opus 5.5 is the default model.
- You need outputs longer than 65.5K tokens: Opus 5.5 writes up to 128K.
- Your prompts exceed 200K input tokens, where Gemini 3.1 Pro Preview's rates rise to $4 input and $18 output while Opus 5.5 stays flat up to 1M.
- You use GitHub Copilot, which offers Opus 5.5 and retired Gemini 3.1 Pro Preview on September 1, 2026.
Choose Gemini 3.1 Pro Preview if
- You want the lower cost: $2 input and $12 output per million tokens, against $4 and $20.
- You use Gemini CLI, where it is the Pro half of the default auto model.
- Your agent depends on custom tools, which Google says the customtools endpoint prioritizes better alongside bash.
- Your sessions write the cache often, since Google bills written tokens as ordinary input.
Side by side
Specs and prices
| Fact | Claude Opus 5.5 | Gemini 3.1 Pro Preview |
|---|---|---|
| Maker | Anthropic | |
| API model id | claude-opus-5-5 | gemini-3.1-pro-preview |
| Released | September 22, 2026 | February 19, 2026 |
| Status | Current | Preview |
| Context window | 1M tokens | 1.05M tokens |
| Max output | 128K tokens | 65.5K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $4 | $2 |
| Cache hit, per 1M | $0.20 | $0.20 |
| Cache write, per 1M | $5 (5-minute), $8 (1-hour) | $2 (same as input) |
| Output, per 1M tokens | $20 | $12 |
| Runs in | Claude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot | Cursor, Gemini CLI, OpenCode, and OpenRouter |
Standard API rates in US dollars, as published by each maker (Claude Opus 5.5: September 26, 2026; Gemini 3.1 Pro Preview: September 28, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Claude Opus 5.5: The full 1M context window is billed at standard rates. Fast mode, a research preview on the Claude API, costs $8 input and $40 output per million tokens. Gemini 3.1 Pro Preview: Prompts over 200K input tokens cost $4 input, $0.40 cached, and $18 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Claude Opus 5.5 | Gemini 3.1 Pro Preview |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $4.40 | $2.00 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.80 | $0.42 |
| Output-heavy generation, 30K input, 80K output | $1.72 | $1.02 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $484.00 | $220.00 |
| Where the session’s cost goes | ||
| Cache writes | $2.60 | $0.80 |
| Cache reads | $0.40 | $0.40 |
| Uncached input | $0.40 | $0.20 |
| Output | $1.00 | $0.60 |
- caching saves on the session with Claude Opus 5.5 (60%)
- $6.60
- caching saves on the session with Gemini 3.1 Pro Preview (64%)
- $3.60
Equal cache hits, unequal cache writes
Claude Opus 5.5 and Gemini 3.1 Pro Preview arrive at the same cache hit price by different routes. Anthropic discounts an Opus 5.5 hit to 5% of its $4 input price, and Google discounts a Gemini hit to 10% of its $2 input price. Both land on $0.20 per million, so the session's 2M cached tokens cost $0.40 on each.
Writes are a different story. Anthropic charges $5 per million for a 5-minute write on Opus 5.5 and $8 for a 1-hour write. Google publishes no separate write price, so written tokens cost the $2 input rate. Cache writes total $2.60 on Opus 5.5 and $0.80 on Gemini, and that $1.80 is most of the $2.40 session difference.
The rest comes from input, $0.40 against $0.20, and output, $1.00 against $0.60. A month of 110 sessions comes to $484.00 on Opus 5.5 and $220.00 on Gemini, API-equivalent estimates that a Claude subscription or Gemini plan would price differently.
Why output-heavy work shows the smallest gap
Apart from cache hits, output is where the two are closest on price: $20 against $12 per million, a 1.7x gap, compared with 2x on input. The output-heavy generation example reflects that, at $1.72 against $1.02. The uncached review lands between the two at 1.9x, $0.80 against $0.42.
Thinking settings affect this more than list prices do. Gemini 3.1 Pro Preview's thinking cannot be turned off and defaults to high, while Opus 5.5 defaults to medium effort. The example holds output at a fixed token count on both, so a model that thinks longer on your tasks will cost more than the table shows. Opus 5.5's tokenizer also counts about 30% more tokens than earlier Claude models for the same text.
Preview status, long prompts, and tool support
Gemini 3.1 Pro Preview has been in preview since February 2026, and Google has announced Gemini 3.5 Pro, which is not yet released. It stays in Gemini CLI as the Pro half of the default auto model, and in Cursor, OpenCode, and OpenRouter. Gemini API key users of the CLI get its customtools endpoint at the same price.
For prompts over 200K input tokens, Google charges $4 input, $0.40 cached, and $18 output per million. Opus 5.5 bills its whole 1M window at the standard $4 and $20. Gemini also caps output at 65.5K tokens per response, about half of Opus 5.5's 128K. Explicit Gemini caches add a storage charge of $4.50 per million tokens per hour on Pro models, which the estimates here leave out.
Anthropic describes Opus 5.5 as built for long-running agentic coding and names long, sprawling jobs such as codebase-wide migrations as a particular strength. Google positions Gemini 3.1 Pro Preview for deep reasoning and agentic coding, and for software engineering that needs precise tool use and reliable multi-step execution.
Prompt caching
How each maker bills cached tokens
Anthropic
Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.
A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.
Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.
In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.
Source: Anthropic: Prompt caching
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what Claude Opus 5.5 and Gemini 3.1 Pro Preview really cost you.
everyaitoken reads your Claude Code, Cursor, OpenCode, OpenRouter, and Gemini CLI history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is Gemini 3.1 Pro Preview cheaper than Claude Opus 5.5?
Yes, on every workload here. The example agentic session costs $2.00 against $4.40, the uncached review $0.42 against $0.80, and output-heavy generation $1.02 against $1.72.
Why do both models charge $0.20 for a cache hit?
Anthropic prices an Opus 5.5 cache hit at 0.05x input instead of its usual 0.1x, and Opus 5.5's input costs twice Gemini's. Google charges 10% of input for a hit on Gemini 3.1 Pro Preview. The two discounts happen to meet at the same price.
What happens above 200K input tokens?
Gemini 3.1 Pro Preview's rates rise to $4 input, $0.40 cached, and $18 output per million tokens. Opus 5.5 has no long-context tier, so its full 1M window bills at $4 input and $20 output.
How can I compare the two on my own work?
EveryToken reads your Claude Code and Gemini CLI history on a Mac and prices each request at the makers' API rates, per model, including what caching saved or cost. That shows whether thinking length or cache writes drive the gap in practice.
Sources
- Anthropic: Pricing
- Anthropic docs: Claude Opus 5.5
- Anthropic: Claude Opus 5.5
- Anthropic docs: Fast mode
- Claude Code docs: Model configuration
- Cursor docs: Claude Opus 5.5
- OpenRouter: Claude Opus 5.5
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Google: Gemini API pricing
- Google docs: Gemini 3.1 Pro Preview
- Google: Gemini models
- Google: Gemini 3.1 Pro
- Google DeepMind: Gemini
- Gemini CLI source: model configuration
- Cursor docs: Gemini 3.1 Pro
- OpenRouter: Gemini 3.1 Pro Preview
- Anthropic: Prompt caching
- Google: Context caching