Model comparison
Claude Fable 5.1 vs Gemini 3.8 Flash: the 14.8x question
A cached coding session costs 14.8x more on Claude Fable 5.1 than on Gemini 3.8 Flash. How Flash's introductory rates, output cap, and free tier factor in.
· Prices as of September 28, 2026
Claude Fable 5.1
Anthropic · Released September 1, 2026
Anthropic's most capable generally available model, aimed at demanding reasoning and long-horizon agentic coding. Anthropic suggests it when Opus-tier results fall short.
Claude Fable 5.1 facts and comparisonsGemini 3.8 Flash
Google · Released September 2, 2026
Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.
Gemini 3.8 Flash facts and comparisons
The short answer
Gemini 3.8 Flash costs a small fraction of Claude Fable 5.1: the example agentic coding session is $0.71 on Flash and $10.50 on Fable 5.1, a 14.8x gap. Google pitches Flash for long-horizon software engineering at Flash prices, while Anthropic suggests Fable 5.1 for demanding work where Opus-tier results fall short. Flash fits high-volume agent loops and Fable 5.1 the hardest tasks, with one caveat: Flash's introductory rates end on December 31, 2026.
Choose Claude Fable 5.1 if
- Opus-tier models have fallen short on the task, which is the case Anthropic names for reaching for Fable 5.1.
- You need responses longer than 65.5K tokens, up to Fable 5.1's 128K.
- You work in Claude Code, where /model fable switches to it, and want Anthropic's most capable generally available model.
Choose Gemini 3.8 Flash if
- You run agent loops at volume: 110 sessions a month cost $78.38 on Flash against $1,155.00 on Fable 5.1.
- You use Gemini CLI with a Gemini API key or Vertex AI, where Flash is the Flash half of the default auto model.
- You want to start on the Gemini API free tier, which covers Flash's input, output, and caching.
- Google's claims match what you need: multi-step planning and tool orchestration with fewer failed loops, and better robustness against prompt injection.
Side by side
Specs and prices
| Fact | Claude Fable 5.1 | Gemini 3.8 Flash |
|---|---|---|
| Maker | Anthropic | |
| API model id | claude-fable-5-1 | gemini-3.8-flash |
| Released | September 1, 2026 | September 2, 2026 |
| Status | Current | Current |
| Context window | 1M tokens | 1.05M tokens |
| Max output | 128K tokens | 65.5K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $10 | $0.75 |
| Cache hit, per 1M | $0.25 | $0.075 |
| Cache write, per 1M | $12.50 (5-minute), $20 (1-hour) | $0.75 (same as input) |
| Output, per 1M tokens | $50 | $3.75 |
| Runs in | Claude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot | Cursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker (Claude Fable 5.1: September 26, 2026; Gemini 3.8 Flash: September 28, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Claude Fable 5.1: The full 1M context window is billed at standard rates. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Claude Fable 5.1 | Gemini 3.8 Flash |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $10.50 | $0.71 |
| Large one-off review, 150K input with no cache hits, 10K output | $2.00 | $0.15 |
| Output-heavy generation, 30K input, 80K output | $4.30 | $0.32 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $1,155.00 | $78.38 |
| Where the session’s cost goes | ||
| Cache writes | $6.50 | $0.30 |
| Cache reads | $0.50 | $0.15 |
| Uncached input | $1.00 | $0.08 |
| Output | $2.50 | $0.19 |
- caching saves on the session with Claude Fable 5.1 (62%)
- $17.00
- caching saves on the session with Gemini 3.8 Flash (66%)
- $1.35
What a 14.8x gap adds up to in a month
At list prices, Claude Fable 5.1 charges 13.3x what Gemini 3.8 Flash does for both input and output: $10 against $0.75, and $50 against $3.75 per million tokens. The uncached review and the output-heavy generation stay close to that ratio, at $2.00 against $0.15 and $4.30 against $0.32.
The agentic coding session stretches the gap to 14.8x, $10.50 against $0.71, and cache writes explain the stretch. Anthropic's writes cost 1.25x or 2x input, while Google bills written tokens as plain input, so writes come to $6.50 on Fable 5.1 and $0.30 on Flash. That $6.20 is most of the $9.79 difference.
Reads work the other way. Anthropic's 97.5% discount on a Fable 5.1 hit is deeper than Google's 90%, so the per-token hit gap is only 3.3x, $0.25 against $0.075. Over 110 sessions a month the totals are $1,155.00 and $78.38, a difference of $1,076.62. These are API-equivalent estimates at published rates, not subscription prices.
Flash's introductory pricing ends December 31, 2026
Google lists Flash's current rates as introductory through December 31, 2026. From January 1, 2027 it charges $1.50 input, $0.15 cached, and $7.50 output per million tokens, exactly twice today's rates. Every Flash cost on this page doubles with it, and the gap to Fable 5.1 roughly halves.
A budget that runs into 2027 should use the higher Flash rates. Explicit caching on Flash models also carries a storage charge of $0.50 to $1 per million tokens per hour, which these estimates leave out.
Output limits and thinking settings
Both models take roughly a million tokens of context, 1M on Fable 5.1 and 1.05M on Flash. The output ceilings differ more: Flash writes at most 65.5K tokens per response, and Fable 5.1 up to 128K.
Thinking defaults change how many tokens each writes. Flash defaults to medium thinking, but Gemini CLI sends high. Fable 5.1 always runs adaptive thinking, with effort from low to max and high as the default. The tables hold output fixed, so a model that thinks longer on your tasks will cost more than they show, and Fable 5.1's newer tokenizer counts about 30% more tokens than earlier Claude models for the same text.
Anthropic describes Fable 5.1's strengths as long-running agentic coding and multistep research. Google calls Gemini 3.8 Flash "Our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows." If you send routine loops to one and hard problems to the other, remember that switching restarts the cache, since each model keeps its own.
Prompt caching
How each maker bills cached tokens
Anthropic
Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.
A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.
Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.
In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.
Source: Anthropic: Prompt caching
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what Claude Fable 5.1 and Gemini 3.8 Flash really cost you.
everyaitoken reads your Claude Code, Cursor, OpenCode, OpenRouter, and Gemini CLI history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
How much cheaper is Gemini 3.8 Flash than Claude Fable 5.1?
At current rates, Flash's input and output cost 93% less. The example agentic session is $0.71 on Flash against $10.50 on Fable 5.1, and a month of 110 sessions $78.38 against $1,155.00.
Will Gemini 3.8 Flash get more expensive?
Yes. Its introductory rates run through December 31, 2026. From January 1, 2027, Google charges $1.50 input, $0.15 cached, and $7.50 output per million tokens.
Is Gemini 3.8 Flash free to use?
The Gemini API free tier covers its input, output, and caching. Paid use costs $0.75 input and $3.75 output per million tokens until the end of 2026, then twice that.
How can I compare my own spend on both?
EveryToken reads your local Claude Code and Gemini CLI history on a Mac and prices each request at Anthropic's and Google's API rates, split by model. It also shows what caching saved or cost, which is where these two models differ most.
Sources
- Anthropic: Pricing
- Anthropic docs: Claude Fable 5.1
- Anthropic: Claude Fable 5.1 and Claude Mythos 5.1
- Claude Code docs: Model configuration
- Cursor docs: Claude Fable 5.1
- OpenRouter: Claude Fable 5.1
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Google: Gemini API pricing
- Google docs: Gemini 3.8 Flash
- Google: Gemini models
- Google: Latest Gemini model
- Google DeepMind: Gemini 3.8 Flash model card
- Google: Gemini 3.8 Flash
- Gemini CLI source: model configuration
- Cursor docs: Gemini 3.8 Flash
- OpenRouter: Gemini 3.8 Flash
- Anthropic: Prompt caching
- Google: Context caching