Skip to content

Model comparison

Gemini 3.8 Flash or Claude Opus 5.5 for agentic coding?

Claude Opus 5.5 costs 6.2x as much as Gemini 3.8 Flash on a cached coding session at Flash's introductory rates. What changes in 2027, and where each fits.

· Prices as of September 28, 2026

  • Claude Opus 5.5

    Anthropic · Released September 22, 2026

    Anthropic's recommended starting model for most work, built for long-running agentic coding. Anthropic says it matches Claude Fable 5.1 on most work at a much lower price.

    Claude Opus 5.5 facts and comparisons
  • Gemini 3.8 Flash

    Google · Released September 2, 2026

    Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.

    Gemini 3.8 Flash facts and comparisons

The short answer

Gemini 3.8 Flash is far cheaper: the example agentic coding session costs $0.71 on it and $4.40 on Claude Opus 5.5, a 6.2x gap, at introductory Flash rates that run through December 31, 2026. Opus 5.5 is Anthropic's recommended starting model and Claude Code's default, while Google aims Gemini 3.8 Flash at long-horizon software engineering and agents at Flash prices. Choose Flash for high-volume agent loops in Gemini CLI, and Opus 5.5 for Claude Code work or outputs longer than 65.5K tokens.

Choose Claude Opus 5.5 if

  • Claude Code is your main tool, and Opus 5.5 is already its default model.
  • You need up to 128K output tokens in a single response.
  • You run long, sprawling jobs such as codebase-wide migrations and audits, a strength Anthropic names for Opus 5.5.

Choose Gemini 3.8 Flash if

  • You run many agent sessions: 110 a month cost $78.38 on Flash and $484.00 on Opus 5.5 at current rates.
  • You use Gemini CLI with a Gemini API key or Vertex AI, where Flash is half of the default auto model.
  • You want to start free, since the Gemini API free tier covers Flash's input, output, and caching.
  • Google's claims match your agents: multi-step planning and tool orchestration with fewer failed loops, and better robustness against prompt injection.

Side by side

Specs and prices

FactClaude Opus 5.5Gemini 3.8 Flash
MakerAnthropicGoogle
API model idclaude-opus-5-5gemini-3.8-flash
ReleasedSeptember 22, 2026September 2, 2026
StatusCurrentCurrent
Context window1M tokens1.05M tokens
Max output128K tokens65.5K tokens
Open weightsNoNo
Input, per 1M tokens$4$0.75
Cache hit, per 1M$0.20$0.075
Cache write, per 1M$5 (5-minute), $8 (1-hour)$0.75 (same as input)
Output, per 1M tokens$20$3.75
Runs inClaude Code, Cursor, OpenCode, OpenRouter, and GitHub CopilotCursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker (Claude Opus 5.5: September 26, 2026; Gemini 3.8 Flash: September 28, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Claude Opus 5.5: The full 1M context window is billed at standard rates. Fast mode, a research preview on the Claude API, costs $8 input and $40 output per million tokens. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadClaude Opus 5.5Gemini 3.8 Flash
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$4.40$0.71
Large one-off review, 150K input with no cache hits, 10K output$0.80$0.15
Output-heavy generation, 30K input, 80K output$1.72$0.32
A month of sessions, 110 sessions: 5 a day, 22 working days$484.00$78.38
Where the session’s cost goes
Cache writes$2.60$0.30
Cache reads$0.40$0.15
Uncached input$0.40$0.08
Output$1.00$0.19
caching saves on the session with Claude Opus 5.5 (60%)
$6.60
caching saves on the session with Gemini 3.8 Flash (66%)
$1.35

How a 5.3x price list becomes a 6.2x session

Claude Opus 5.5 charges 5.3x the list price of Gemini 3.8 Flash for input and output: $4 against $0.75, and $20 against $3.75 per million tokens. The uncached review, $0.80 against $0.15, and the output-heavy generation, $1.72 against $0.32, stay close to that ratio.

The agentic session is where the ratio grows. Anthropic's 5-minute write costs $5 and its 1-hour write $8 per million on Opus 5.5, while Google bills Flash's written tokens as ordinary input, $0.75. Cache writes reach $2.60 on Opus 5.5 against $0.30 on Flash, $2.30 of the $3.69 difference.

Cache reads narrow things slightly. Anthropic's 95% discount on an Opus 5.5 hit is deeper than Google's 90%, so hits are only 2.7x apart, $0.20 against $0.075. The session ends at $4.40 against $0.71, and 110 sessions a month at $484.00 against $78.38, a $405.62 difference in API-equivalent terms.

What January 1, 2027 changes

Google lists Flash's prices as introductory rates through December 31, 2026. From January 1, 2027, Flash costs $1.50 input, $0.15 cached, and $7.50 output per million tokens, double the current rates. Every Flash workload cost doubles with them, so the gap to Opus 5.5 roughly halves.

Opus 5.5's published price notes mention no scheduled change. If you are picking a model for a project that runs into 2027, compare against the higher Flash rates rather than the introductory ones.

Claude Code's default against Gemini CLI's Flash

Each model is a default in its maker's command-line tool. Opus 5.5 is Claude Code's default model on Pro, Max, Team, and Enterprise plans and with an Anthropic API key. Gemini 3.8 Flash is the Flash half of Gemini CLI's default auto model for Gemini API key and Vertex AI users, and it is also available in Cursor, OpenCode, OpenRouter, and GitHub Copilot.

They also differ in how much they write per response. Flash caps output at 65.5K tokens and Opus 5.5 at 128K. Flash defaults to medium thinking, and Gemini CLI raises it to high. Opus 5.5 defaults to medium effort, and Anthropic says changing effort keeps the prompt cache. Both accept about a million tokens of context, 1M on Opus 5.5 and 1.05M on Flash.

Anthropic claims Opus 5.5 produces output more than 30% faster than Claude Opus 5 and says it "costs 40% less to run than Opus 5." Google calls Gemini 3.8 Flash its most intelligent Flash model. Neither claim has been checked for this page, and the makers' tokenizers differ, so the same task can use different token counts on each.

Prompt caching

How each maker bills cached tokens

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

Google

Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.

Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.

On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.

Source: Google: Context caching

Your own numbers

See what Claude Opus 5.5 and Gemini 3.8 Flash really cost you.

everyaitoken reads your Claude Code, Cursor, OpenCode, OpenRouter, and Gemini CLI history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

How much cheaper is Gemini 3.8 Flash than Claude Opus 5.5?

At current rates, input and output cost 81% less on Flash. The example agentic session costs $0.71 against $4.40, which is 84% less. From January 1, 2027, Flash's prices double.

When do Gemini 3.8 Flash's introductory prices end?

They end on December 31, 2026. After that, Flash's rates become $1.50 per million input tokens, $0.15 per million cached tokens, and $7.50 per million output tokens.

Which model has the larger output limit?

Claude Opus 5.5, at 128K tokens per response. Gemini 3.8 Flash writes at most 65.5K, which can matter for large generated files or long plans.

Can I track both models in one place?

EveryToken reads your local Claude Code and Gemini CLI history on a Mac, prices each request at API rates, and splits the cost and cache savings by model. It is a $9 one-time purchase.

  • Claude Fable 5.1 vs Claude Opus 5.5

    Claude Fable 5.1 lists at 2.5x the rates of Claude Opus 5.5. What that premium buys, why cheap cache hits barely narrow it, and when Anthropic suggests it.

  • Claude Fable 5.1 vs Gemini 3.8 Flash

    A cached coding session costs 14.8x more on Claude Fable 5.1 than on Gemini 3.8 Flash. How Flash's introductory rates, output cap, and free tier factor in.

  • Claude Haiku 4.5 vs Gemini 3.8 Flash

    Gemini 3.8 Flash undercuts Claude Haiku 4.5 on introductory rates, $0.71 against $1.20 per cached session. From January 1, 2027, its rates top Haiku's.

  • Claude Opus 5.5 vs Claude Opus 4.8

    Claude Opus 5.5 lists 20% below Claude Opus 4.8 and bills cache hits at $0.20, not $0.50. What staying costs, and when Anthropic still recommends Opus 4.8.

  • Claude Opus 5.5 vs Claude Haiku 4.5

    Claude Opus 5.5 lists at 4x the rates of Claude Haiku 4.5, yet a cached coding session costs 3.7x. Context and output limits, and Haiku 4.5's retirement date.

  • Claude Opus 5.5 vs Claude Opus 5

    Claude Opus 5.5 is 20% cheaper per token than Claude Opus 5, and cache hits cost 60% less. What that means for Claude Code sessions, and what stays the same.