Skip to content

Model comparison

Claude Haiku 4.5 vs Gemini 3.8 Flash: cost now and in 2027

Gemini 3.8 Flash undercuts Claude Haiku 4.5 on introductory rates, $0.71 against $1.20 per cached session. From January 1, 2027, its rates top Haiku's.

· Prices as of September 28, 2026

  • Claude Haiku 4.5

    Anthropic · Released October 15, 2025

    Anthropic's fastest and cheapest current model, aimed at latency-sensitive work and at coding sub-agents.

    Claude Haiku 4.5 facts and comparisons
  • Gemini 3.8 Flash

    Google · Released September 2, 2026

    Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.

    Gemini 3.8 Flash facts and comparisons

The short answer

Through December 31, 2026, Gemini 3.8 Flash is the cheaper model: the example agentic coding session costs $0.71 against $1.20 on Claude Haiku 4.5, a 1.7x gap. From January 1, 2027, Google lists $1.50 input, $0.15 cached, and $7.50 output for Gemini 3.8 Flash, all above Haiku 4.5's rates. Gemini 3.8 Flash also has a 1.05M context window against 200K, while Haiku 4.5 remains Anthropic's small model for Claude Code sub-agents.

Choose Claude Haiku 4.5 if

  • You run sub-agents in Claude Code, the use Anthropic names for Haiku 4.5 alongside latency-sensitive work.
  • You are planning costs for 2027, when Gemini 3.8 Flash's listed rates rise above Haiku 4.5's.
  • You want a fixed thinking budget: Haiku 4.5 uses manual extended thinking rather than thinking levels.

Choose Gemini 3.8 Flash if

  • You want the lower cost through December 31, 2026: $0.75 input and $3.75 output per million against $1 and $5.
  • Your tasks need more than 200K tokens of context, since Gemini 3.8 Flash accepts 1.05M.
  • You use Gemini CLI with an API key or Vertex AI, where it already serves as the Flash half of the default auto model.
  • You want to try it free first: the Gemini API free tier covers its input, output, and caching.

Side by side

Specs and prices

FactClaude Haiku 4.5Gemini 3.8 Flash
MakerAnthropicGoogle
API model idclaude-haiku-4-5gemini-3.8-flash
ReleasedOctober 15, 2025September 2, 2026
StatusCurrentCurrent
Context window200K tokens1.05M tokens
Max output64K tokens65.5K tokens
Open weightsNoNo
Input, per 1M tokens$1$0.75
Cache hit, per 1M$0.10$0.075
Cache write, per 1M$1.25 (5-minute), $2 (1-hour)$0.75 (same as input)
Output, per 1M tokens$5$3.75
Runs inClaude Code, Cursor, OpenCode, OpenRouter, and GitHub CopilotCursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker (Claude Haiku 4.5: September 26, 2026; Gemini 3.8 Flash: September 28, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadClaude Haiku 4.5Gemini 3.8 Flash
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$1.20$0.71
Large one-off review, 150K input with no cache hits, 10K output$0.20$0.15
Output-heavy generation, 30K input, 80K output$0.43$0.32
A month of sessions, 110 sessions: 5 a day, 22 working days$132.00$78.38
Where the session’s cost goes
Cache writes$0.65$0.30
Cache reads$0.20$0.15
Uncached input$0.10$0.08
Output$0.25$0.19
caching saves on the session with Claude Haiku 4.5 (56%)
$1.55
caching saves on the session with Gemini 3.8 Flash (66%)
$1.35

Why Gemini 3.8 Flash costs less through 2026

On today's introductory rates, Gemini 3.8 Flash is 25% cheaper than Claude Haiku 4.5 on input, output, and cache hits: $0.75 input, $3.75 output, and $0.075 per cache hit, per million tokens, against $1 input, $5 output, and $0.10 per hit. The large one-off review costs $0.15 against $0.20, and the output-heavy generation $0.32 against $0.43.

The cached session gap is bigger, 1.7x, because the two makers bill cache writes differently. Google bills written tokens as ordinary input. Anthropic charges 1.25x input for a 5-minute write and 2x for a 1-hour write, which is $2 per million on Haiku 4.5. Writes cost $0.65 on Haiku 4.5 against $0.30 on Gemini 3.8 Flash, $0.35 of the $0.49 difference.

Across 110 sessions a month that is $132.00 against $78.38. Caching saves 66% of the uncached session cost on Gemini 3.8 Flash and 56% on Haiku 4.5.

Which is cheaper after January 1, 2027?

Gemini 3.8 Flash's current rates are introductory. From January 1, 2027, Google lists $1.50 per million input tokens, $0.15 per million cached tokens, and $7.50 per million output tokens. Each of those is above Haiku 4.5's list rate, and with written tokens billed at the new $1.50 input rate, the example session would then cost more on Gemini 3.8 Flash than on Haiku 4.5 as well.

That comparison may not be the one you face in 2027. Haiku 4.5's retirement is listed for no sooner than October 15, 2026, and Anthropic announced Claude Haiku 5.5 on September 22, 2026. Its pricing is not in the data behind this post. For planning past the new year, treat the current gap as temporary on both sides.

How Google and Anthropic position these two

Google calls Gemini 3.8 Flash its most intelligent Flash model, built for long-horizon software engineering, autonomous agents, and complex enterprise workflows. It claims deterministic tool execution, multi-file refactoring, fewer failed agent loops, and better robustness against prompt injection. It thinks at a medium level by default, and Gemini CLI asks for high.

Anthropic aims Haiku 4.5 at latency-sensitive work and coding sub-agents, and says its coding performance is similar to Claude Sonnet 4's while running more than twice as fast. It uses Anthropic's older tokenizer, which counts fewer tokens for the same text than Claude models from 4.7 on.

The limits separate them most. Haiku 4.5 offers a 200K context window and up to 64K tokens of output. Gemini 3.8 Flash has 1.05M and 65.5K. Both run in Cursor, OpenRouter, OpenCode, and GitHub Copilot, with Haiku 4.5 in Claude Code and Gemini 3.8 Flash in Gemini CLI.

Prompt caching

How each maker bills cached tokens

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

Google

Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.

Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.

On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.

Source: Google: Context caching

Your own numbers

See what Claude Haiku 4.5 and Gemini 3.8 Flash really cost you.

everyaitoken reads your Claude Code, Cursor, OpenCode, OpenRouter, and Gemini CLI history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is Gemini 3.8 Flash cheaper than Claude Haiku 4.5?

Until December 31, 2026, yes: 25% less per token and $0.71 against $1.20 on the example cached session. From January 1, 2027, its listed rates of $1.50 input and $7.50 output are higher than Haiku 4.5's $1 and $5.

When does Gemini 3.8 Flash's introductory pricing end?

The introductory rates run through December 31, 2026. Google lists the new rates from January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.

Is Claude Haiku 4.5 being replaced?

Anthropic announced Claude Haiku 5.5 on September 22, 2026. Haiku 4.5 retires no sooner than October 15, 2026.

How can I see my own cost on each model?

EveryToken prices every request in your local Claude Code and Gemini CLI history at API rates and splits the total by model, cache savings included. The totals are API-equivalent estimates, so a subscription or free tier will charge differently.

  • Claude Fable 5.1 vs Gemini 3.8 Flash

    A cached coding session costs 14.8x more on Claude Fable 5.1 than on Gemini 3.8 Flash. How Flash's introductory rates, output cap, and free tier factor in.

  • Claude Haiku 4.5 vs Gemini 3.5 Flash-Lite

    Claude Haiku 4.5 and Gemini 3.5 Flash-Lite both target sub-agent work. Haiku 4.5 costs 3.5x as much per cached session, but only 2x as much per output token.

  • Claude Opus 5.5 vs Claude Haiku 4.5

    Claude Opus 5.5 lists at 4x the rates of Claude Haiku 4.5, yet a cached coding session costs 3.7x. Context and output limits, and Haiku 4.5's retirement date.

  • Claude Sonnet 5 vs Claude Haiku 4.5

    Claude Haiku 4.5 costs half as much as Claude Sonnet 5, with a 200K context window and a retirement date ahead. When the cheaper model fits a coding workflow.

  • Claude Sonnet 5 vs Gemini 3.8 Flash

    Gemini 3.8 Flash runs a cached coding session for $0.71 against $2.40 on Claude Sonnet 5, on introductory rates that end December 31, 2026.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash

    Gemini 3.8 Flash costs about half of Gemini 3.5 Flash on introductory rates. What changes on January 1, 2027, and which one Gemini CLI picks for you.