Model comparison
Claude Fable 5.1 vs Gemini 3.1 Pro Preview: 5.3x per session
A cached coding session costs 5.3x more on Claude Fable 5.1 than on Gemini 3.1 Pro Preview, mostly from cache writes. Output limits and preview status compared.
· Prices as of September 28, 2026
Claude Fable 5.1
Anthropic · Released September 1, 2026
Anthropic's most capable generally available model, aimed at demanding reasoning and long-horizon agentic coding. Anthropic suggests it when Opus-tier results fall short.
Claude Fable 5.1 facts and comparisonsGemini 3.1 Pro Preview
Google · Released February 19, 2026 · Preview
Google's only current Pro model, still in preview, positioned for deep reasoning and agentic coding.
Gemini 3.1 Pro Preview facts and comparisons
The short answer
Gemini 3.1 Pro Preview is the much cheaper model: the example agentic coding session costs $2.00 on it and $10.50 on Claude Fable 5.1, a 5.3x gap. Most of the difference is cache writes, which Google bills as ordinary input while Anthropic adds a premium. Fable 5.1 suits work that needs long outputs, flat long-context pricing, or a generally available model, and Gemini 3.1 Pro Preview suits Gemini CLI users who can accept a preview.
Choose Claude Fable 5.1 if
- You need long responses: Fable 5.1 writes up to 128K tokens, against 65.5K on Gemini 3.1 Pro Preview.
- Your prompts often pass 200K input tokens, where Gemini 3.1 Pro Preview moves to $4 input and $18 output while Fable 5.1 bills its whole 1M window at standard rates.
- You want a generally available model in Claude Code, Cursor, or GitHub Copilot rather than a preview.
- Your tasks fit Anthropic's positioning for Fable 5.1: demanding reasoning and long-horizon agentic coding.
Choose Gemini 3.1 Pro Preview if
- Cost comes first: the session costs $2.00 instead of $10.50, and 110 of them $220.00 instead of $1,155.00.
- You code in Gemini CLI, whose default auto model uses Gemini 3.1 Pro Preview as its Pro half.
- Your agent leans on custom tools, which Google says its separate customtools endpoint prioritizes better alongside bash.
Side by side
Specs and prices
| Fact | Claude Fable 5.1 | Gemini 3.1 Pro Preview |
|---|---|---|
| Maker | Anthropic | |
| API model id | claude-fable-5-1 | gemini-3.1-pro-preview |
| Released | September 1, 2026 | February 19, 2026 |
| Status | Current | Preview |
| Context window | 1M tokens | 1.05M tokens |
| Max output | 128K tokens | 65.5K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $10 | $2 |
| Cache hit, per 1M | $0.25 | $0.20 |
| Cache write, per 1M | $12.50 (5-minute), $20 (1-hour) | $2 (same as input) |
| Output, per 1M tokens | $50 | $12 |
| Runs in | Claude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot | Cursor, Gemini CLI, OpenCode, and OpenRouter |
Standard API rates in US dollars, as published by each maker (Claude Fable 5.1: September 26, 2026; Gemini 3.1 Pro Preview: September 28, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Claude Fable 5.1: The full 1M context window is billed at standard rates. Gemini 3.1 Pro Preview: Prompts over 200K input tokens cost $4 input, $0.40 cached, and $18 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Claude Fable 5.1 | Gemini 3.1 Pro Preview |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $10.50 | $2.00 |
| Large one-off review, 150K input with no cache hits, 10K output | $2.00 | $0.42 |
| Output-heavy generation, 30K input, 80K output | $4.30 | $1.02 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $1,155.00 | $220.00 |
| Where the session’s cost goes | ||
| Cache writes | $6.50 | $0.80 |
| Cache reads | $0.50 | $0.40 |
| Uncached input | $1.00 | $0.20 |
| Output | $2.50 | $0.60 |
- caching saves on the session with Claude Fable 5.1 (62%)
- $17.00
- caching saves on the session with Gemini 3.1 Pro Preview (64%)
- $3.60
Where the 5.3x session gap comes from
On list prices, Claude Fable 5.1 charges 5x the input price of Gemini 3.1 Pro Preview, $10 against $2 per million tokens, and 4.2x its output price, $50 against $12. The example agentic coding session opens a wider gap than either, 5.3x, because of how each maker bills cache writes.
Anthropic charges 1.25x input for a 5-minute write and 2x for a 1-hour write, which is $12.50 and $20 per million on Fable 5.1. Google publishes no separate write price, so written tokens cost ordinary input, $2. In the session, cache writes cost $6.50 on Fable 5.1 and $0.80 on Gemini, and that $5.70 accounts for most of the $8.50 difference.
Cache reads, by contrast, are nearly level. Fable 5.1's hit costs 2.5% of its input price, $0.25 per million, against Gemini's $0.20, so 2M cached tokens cost $0.50 and $0.40. Without cache writes in the mix the gap shrinks: 4.8x on the uncached review and 4.2x on output-heavy generation, where Gemini's output price weighs most.
Output ceiling and the 200K price line
Gemini 3.1 Pro Preview writes at most 65.5K tokens per response. Fable 5.1 writes up to 128K, which matters for large generated files, long plans, or extended reasoning that ends in a big patch.
The long-context rules differ too. Google raises Gemini 3.1 Pro Preview's rates for prompts over 200K input tokens to $4 input, $0.40 cached, and $18 output per million. Fable 5.1 has no long-context tier, so every token in its 1M window costs the same. Each request in the example session stays under 200K, so the cost table uses Gemini's lower tier, but an agent that sends whole repositories in one prompt would pay the higher one.
A preview model against a generally available one
Gemini 3.1 Pro Preview has been in preview since February 19, 2026. Google has announced a successor, Gemini 3.5 Pro, that is not out yet, and GitHub Copilot stopped offering Gemini 3.1 Pro Preview on September 1, 2026. It remains in Gemini CLI as the Pro half of the default auto model, where Gemini API key users get the customtools endpoint at the same price, and in Cursor, OpenCode, and OpenRouter.
Google describes it as offering "Advanced intelligence, complex problem-solving skills, and powerful agentic and vibe coding capabilities." Its thinking cannot be turned off and defaults to high, which adds output tokens the tables do not model. Fable 5.1 also defaults to high effort, and its tokenizer counts about 30% more tokens than earlier Claude models, so the same prompt yields different counts on each.
Fable 5.1 is available in Claude Code through /model fable, and in Cursor, OpenCode, OpenRouter, and GitHub Copilot. It requires 30-day data retention, so organizations that need zero data retention cannot use it.
Prompt caching
How each maker bills cached tokens
Anthropic
Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.
A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.
Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.
In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.
Source: Anthropic: Prompt caching
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what Claude Fable 5.1 and Gemini 3.1 Pro Preview really cost you.
everyaitoken reads your Claude Code, Cursor, OpenCode, OpenRouter, and Gemini CLI history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Why is Claude Fable 5.1 so much more expensive than Gemini 3.1 Pro Preview?
Its list prices are higher, 5x on input and 4.2x on output, and Anthropic charges a premium to write the cache while Google does not. On the example session, cache writes alone cost $6.50 on Fable 5.1 against $0.80 on Gemini. A month of 110 sessions comes to $1,155.00 against $220.00, as API-equivalent estimates.
Which of the two can I still use in GitHub Copilot?
Only Claude Fable 5.1. Copilot dropped Gemini 3.1 Pro Preview on September 1, 2026, while Fable 5.1 remains on its model list.
Does Gemini's context caching cost extra?
Implicit caching is on by default and applies the discount automatically, though a hit is not assured. Explicit caching assures the discount for a named cache and adds a storage charge of $4.50 per million tokens per hour on Pro models. The estimates here price hits at 10% of input and leave storage out.
Is Gemini 3.5 Pro available yet?
Not yet. Google has announced Gemini 3.5 Pro, but as of September 28, 2026 it is not released, and Gemini 3.1 Pro Preview is Google's current Pro model. EveryToken can price your Gemini CLI and Claude Code sessions on today's models, which gives you a baseline to compare against once a successor arrives.
Sources
- Anthropic: Pricing
- Anthropic docs: Claude Fable 5.1
- Anthropic: Claude Fable 5.1 and Claude Mythos 5.1
- Claude Code docs: Model configuration
- Cursor docs: Claude Fable 5.1
- OpenRouter: Claude Fable 5.1
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Google: Gemini API pricing
- Google docs: Gemini 3.1 Pro Preview
- Google: Gemini models
- Google: Gemini 3.1 Pro
- Google DeepMind: Gemini
- Gemini CLI source: model configuration
- Cursor docs: Gemini 3.1 Pro
- OpenRouter: Gemini 3.1 Pro Preview
- Anthropic: Prompt caching
- Google: Context caching