Model comparison
Claude Opus 4.8 vs Gemini 3.1 Pro Preview: cost and limits
Claude Opus 4.8 costs 3x Gemini 3.1 Pro Preview on a cached coding session, $6.00 against $2.00. How long prompts, fast mode, and preview status shift that.
· Prices as of September 28, 2026
Claude Opus 4.8
Anthropic · Released May 28, 2026 · Previous generation
An Opus upgrade over 4.7 focused on judgment and collaboration. Anthropic still recommends it for cybersecurity work that needs reduced guardrails.
Claude Opus 4.8 facts and comparisonsGemini 3.1 Pro Preview
Google · Released February 19, 2026 · Preview
Google's only current Pro model, still in preview, positioned for deep reasoning and agentic coding.
Gemini 3.1 Pro Preview facts and comparisons
The short answer
Gemini 3.1 Pro Preview costs far less: the example agentic coding session is $2.00 against $6.00 on Claude Opus 4.8, a 3x gap, and a month of sessions $220.00 against $660.00. Claude Opus 4.8 is now a legacy model that Anthropic still recommends for cybersecurity work needing reduced guardrails, while Gemini 3.1 Pro Preview is Google's current Pro model, still in preview. The gap narrows on prompts over 200K input tokens, where Gemini's rates rise and Opus 4.8 keeps its standard rates across its 1M window.
Choose Claude Opus 4.8 if
- You do security work that needs reduced guardrails, the case where Anthropic still points to Opus 4.8.
- You send prompts well past 200K tokens: Anthropic bills Opus 4.8's full 1M window at standard rates.
- You need up to 128K tokens of output in one response, against Gemini's 65.5K.
- You want fast mode, a research preview at $10 input and $50 output per million that Anthropic says runs at 2.5x speed.
Choose Gemini 3.1 Pro Preview if
- Cost is the priority: $2 input and $12 output per million against $5 and $25.
- You rely on Gemini CLI's default auto model, whose Pro half is Gemini 3.1 Pro Preview.
- Your agent mixes custom tools with bash, and you want the customtools endpoint Google offers at the same price.
- Your sessions cache heavily, since Google bills written tokens as ordinary input with no premium.
Side by side
Specs and prices
| Fact | Claude Opus 4.8 | Gemini 3.1 Pro Preview |
|---|---|---|
| Maker | Anthropic | |
| API model id | claude-opus-4-8 | gemini-3.1-pro-preview |
| Released | May 28, 2026 | February 19, 2026 |
| Status | Previous generation | Preview |
| Context window | 1M tokens | 1.05M tokens |
| Max output | 128K tokens | 65.5K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $5 | $2 |
| Cache hit, per 1M | $0.50 | $0.20 |
| Cache write, per 1M | $6.25 (5-minute), $10 (1-hour) | $2 (same as input) |
| Output, per 1M tokens | $25 | $12 |
| Runs in | Claude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot | Cursor, Gemini CLI, OpenCode, and OpenRouter |
Standard API rates in US dollars, as published by each maker (Claude Opus 4.8: September 26, 2026; Gemini 3.1 Pro Preview: September 28, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Claude Opus 4.8: The full 1M context window is billed at standard rates. Fast mode, a research preview on the Claude API, costs $10 input and $50 output per million tokens. Gemini 3.1 Pro Preview: Prompts over 200K input tokens cost $4 input, $0.40 cached, and $18 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Claude Opus 4.8 | Gemini 3.1 Pro Preview |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $6.00 | $2.00 |
| Large one-off review, 150K input with no cache hits, 10K output | $1.00 | $0.42 |
| Output-heavy generation, 30K input, 80K output | $2.15 | $1.02 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $660.00 | $220.00 |
| Where the session’s cost goes | ||
| Cache writes | $3.25 | $0.80 |
| Cache reads | $1.00 | $0.40 |
| Uncached input | $0.50 | $0.20 |
| Output | $1.25 | $0.60 |
- caching saves on the session with Claude Opus 4.8 (56%)
- $7.75
- caching saves on the session with Gemini 3.1 Pro Preview (64%)
- $3.60
Where the 3x session gap comes from
Claude Opus 4.8 charges $5 per million input tokens, $0.50 per million cache hits, $6.25 for a 5-minute cache write, $10 for a 1-hour write, and $25 per million output tokens. Gemini 3.1 Pro Preview charges $2 for input, $0.20 for hits, and $12 for output, and bills written tokens as ordinary input at $2.
On uncached work the gap is 2.1x to 2.4x. The large one-off review costs $1.00 on Opus 4.8 and $0.42 on Gemini, and the output-heavy generation $2.15 against $1.02.
The cached session widens it to 3x, $6.00 against $2.00. Cache writes explain most of that: Anthropic charges 1.25x input for a 5-minute write and 2x for a 1-hour write, so the session's 400K written tokens cost $3.25 on Opus 4.8 and $0.80 on Gemini. That $2.45 is more than half of the $4.00 difference. Over 110 sessions a month the totals are $660.00 and $220.00.
Long prompts narrow the gap
Gemini's rates change with prompt size. Above 200K input tokens it charges $4 per million for input, $0.40 for cached tokens, and $18 for output. Anthropic bills Opus 4.8's full 1M context window at its standard rates, so above 200K the comparison becomes $5 against $4 for input and $25 against $18 for output.
The example session keeps each request under 200K, so none of its figures include Gemini's higher tier. Coding agents that routinely pack a large share of a repository into one request should price that tier in. Google's explicit caching also adds a storage charge on Pro models, $4.50 per million tokens per hour, which the session figures leave out.
A legacy Opus against a preview Pro
Opus 4.8 is a previous model that stays available. Anthropic positions it as an Opus upgrade focused on judgment and collaboration, with the consistency and autonomy to keep working on long-running tasks, and recommends starting at xhigh effort for coding and agentic work. Higher effort writes more tokens, and output is $1.25 of the Opus 4.8 session. For most new work, Anthropic's recommended starting model is now Claude Opus 5.5.
Gemini 3.1 Pro Preview is Google's current Pro model but still a preview, and Google has announced Gemini 3.5 Pro without releasing it. Google gives no way to switch its thinking off on this preview, and the level starts at high. GitHub Copilot retired it on September 1, 2026, while Opus 4.8 remains there, alongside Claude Code, Cursor, OpenRouter, and OpenCode. Gemini 3.1 Pro Preview runs in Gemini CLI, Cursor, OpenRouter, and OpenCode.
Token counts differ between the two. Opus 4.8 uses Anthropic's newer tokenizer, which counts about 30% more tokens than earlier Claude models for the same text, and Google's tokenizer is different again.
Prompt caching
How each maker bills cached tokens
Anthropic
Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.
A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.
Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.
In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.
Source: Anthropic: Prompt caching
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what Claude Opus 4.8 and Gemini 3.1 Pro Preview really cost you.
everyaitoken reads your Claude Code, Cursor, OpenCode, OpenRouter, and Gemini CLI history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
How much more does Claude Opus 4.8 cost than Gemini 3.1 Pro Preview?
About 2.5x on input and cache hits and 2.1x on output. On the example cached session it is 3x, $6.00 against $2.00, because Anthropic charges a premium for cache writes and Google does not.
What does fast mode cost on Claude Opus 4.8?
Fast mode is a research preview on the Claude API at $10 input and $50 output per million tokens. Anthropic says it runs at 2.5x speed.
Does GitHub Copilot still offer Gemini 3.1 Pro Preview?
No. It left GitHub Copilot on September 1, 2026, and remains in Gemini CLI, Cursor, OpenRouter, and OpenCode.
Should I use Claude Opus 4.8 or Claude Opus 5.5?
Anthropic calls Claude Opus 5.5 its recommended starting model for most work and still recommends Opus 4.8 for cybersecurity work that needs reduced guardrails. EveryToken can show what each Claude model costs you in Claude Code, priced at API rates.
Sources
- Anthropic: Pricing
- Anthropic docs: Claude Opus 4.8
- Anthropic: Introducing Claude Opus 4.8
- Anthropic: Claude Opus
- Claude Code docs: Model configuration
- Cursor docs: Claude Opus 4.8
- OpenRouter: Claude Opus 4.8
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Google: Gemini API pricing
- Google docs: Gemini 3.1 Pro Preview
- Google: Gemini models
- Google: Gemini 3.1 Pro
- Google DeepMind: Gemini
- Gemini CLI source: model configuration
- Cursor docs: Gemini 3.1 Pro
- OpenRouter: Gemini 3.1 Pro Preview
- Anthropic: Prompt caching
- Google: Context caching