Model comparison
Claude Haiku 4.5 vs Gemini 3.8 Flash: cost now and in 2027
Gemini 3.8 Flash undercuts Claude Haiku 4.5 on introductory rates, $0.71 against $1.20 per cached session. From January 1, 2027, its rates top Haiku's.
· Prices as of September 28, 2026
Claude Haiku 4.5
Anthropic · Released October 15, 2025
Anthropic's fastest and cheapest current model, aimed at latency-sensitive work and at coding sub-agents.
Claude Haiku 4.5 facts and comparisonsGemini 3.8 Flash
Google · Released September 2, 2026
Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.
Gemini 3.8 Flash facts and comparisons
The short answer
Through December 31, 2026, Gemini 3.8 Flash is the cheaper model: the example agentic coding session costs $0.71 against $1.20 on Claude Haiku 4.5, a 1.7x gap. From January 1, 2027, Google lists $1.50 input, $0.15 cached, and $7.50 output for Gemini 3.8 Flash, all above Haiku 4.5's rates. Gemini 3.8 Flash also has a 1.05M context window against 200K, while Haiku 4.5 remains Anthropic's small model for Claude Code sub-agents.
Choose Claude Haiku 4.5 if
- You run sub-agents in Claude Code, the use Anthropic names for Haiku 4.5 alongside latency-sensitive work.
- You are planning costs for 2027, when Gemini 3.8 Flash's listed rates rise above Haiku 4.5's.
- You want a fixed thinking budget: Haiku 4.5 uses manual extended thinking rather than thinking levels.
Choose Gemini 3.8 Flash if
- You want the lower cost through December 31, 2026: $0.75 input and $3.75 output per million against $1 and $5.
- Your tasks need more than 200K tokens of context, since Gemini 3.8 Flash accepts 1.05M.
- You use Gemini CLI with an API key or Vertex AI, where it already serves as the Flash half of the default auto model.
- You want to try it free first: the Gemini API free tier covers its input, output, and caching.
Side by side
Specs and prices
| Fact | Claude Haiku 4.5 | Gemini 3.8 Flash |
|---|---|---|
| Maker | Anthropic | |
| API model id | claude-haiku-4-5 | gemini-3.8-flash |
| Released | October 15, 2025 | September 2, 2026 |
| Status | Current | Current |
| Context window | 200K tokens | 1.05M tokens |
| Max output | 64K tokens | 65.5K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $1 | $0.75 |
| Cache hit, per 1M | $0.10 | $0.075 |
| Cache write, per 1M | $1.25 (5-minute), $2 (1-hour) | $0.75 (same as input) |
| Output, per 1M tokens | $5 | $3.75 |
| Runs in | Claude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot | Cursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker (Claude Haiku 4.5: September 26, 2026; Gemini 3.8 Flash: September 28, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Claude Haiku 4.5 | Gemini 3.8 Flash |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $1.20 | $0.71 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.20 | $0.15 |
| Output-heavy generation, 30K input, 80K output | $0.43 | $0.32 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $132.00 | $78.38 |
| Where the session’s cost goes | ||
| Cache writes | $0.65 | $0.30 |
| Cache reads | $0.20 | $0.15 |
| Uncached input | $0.10 | $0.08 |
| Output | $0.25 | $0.19 |
- caching saves on the session with Claude Haiku 4.5 (56%)
- $1.55
- caching saves on the session with Gemini 3.8 Flash (66%)
- $1.35
Why Gemini 3.8 Flash costs less through 2026
On today's introductory rates, Gemini 3.8 Flash is 25% cheaper than Claude Haiku 4.5 on input, output, and cache hits: $0.75 input, $3.75 output, and $0.075 per cache hit, per million tokens, against $1 input, $5 output, and $0.10 per hit. The large one-off review costs $0.15 against $0.20, and the output-heavy generation $0.32 against $0.43.
The cached session gap is bigger, 1.7x, because the two makers bill cache writes differently. Google bills written tokens as ordinary input. Anthropic charges 1.25x input for a 5-minute write and 2x for a 1-hour write, which is $2 per million on Haiku 4.5. Writes cost $0.65 on Haiku 4.5 against $0.30 on Gemini 3.8 Flash, $0.35 of the $0.49 difference.
Across 110 sessions a month that is $132.00 against $78.38. Caching saves 66% of the uncached session cost on Gemini 3.8 Flash and 56% on Haiku 4.5.
Which is cheaper after January 1, 2027?
Gemini 3.8 Flash's current rates are introductory. From January 1, 2027, Google lists $1.50 per million input tokens, $0.15 per million cached tokens, and $7.50 per million output tokens. Each of those is above Haiku 4.5's list rate, and with written tokens billed at the new $1.50 input rate, the example session would then cost more on Gemini 3.8 Flash than on Haiku 4.5 as well.
That comparison may not be the one you face in 2027. Haiku 4.5's retirement is listed for no sooner than October 15, 2026, and Anthropic announced Claude Haiku 5.5 on September 22, 2026. Its pricing is not in the data behind this post. For planning past the new year, treat the current gap as temporary on both sides.
How Google and Anthropic position these two
Google calls Gemini 3.8 Flash its most intelligent Flash model, built for long-horizon software engineering, autonomous agents, and complex enterprise workflows. It claims deterministic tool execution, multi-file refactoring, fewer failed agent loops, and better robustness against prompt injection. It thinks at a medium level by default, and Gemini CLI asks for high.
Anthropic aims Haiku 4.5 at latency-sensitive work and coding sub-agents, and says its coding performance is similar to Claude Sonnet 4's while running more than twice as fast. It uses Anthropic's older tokenizer, which counts fewer tokens for the same text than Claude models from 4.7 on.
The limits separate them most. Haiku 4.5 offers a 200K context window and up to 64K tokens of output. Gemini 3.8 Flash has 1.05M and 65.5K. Both run in Cursor, OpenRouter, OpenCode, and GitHub Copilot, with Haiku 4.5 in Claude Code and Gemini 3.8 Flash in Gemini CLI.
Prompt caching
How each maker bills cached tokens
Anthropic
Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.
A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.
Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.
In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.
Source: Anthropic: Prompt caching
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what Claude Haiku 4.5 and Gemini 3.8 Flash really cost you.
everyaitoken reads your Claude Code, Cursor, OpenCode, OpenRouter, and Gemini CLI history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is Gemini 3.8 Flash cheaper than Claude Haiku 4.5?
Until December 31, 2026, yes: 25% less per token and $0.71 against $1.20 on the example cached session. From January 1, 2027, its listed rates of $1.50 input and $7.50 output are higher than Haiku 4.5's $1 and $5.
When does Gemini 3.8 Flash's introductory pricing end?
The introductory rates run through December 31, 2026. Google lists the new rates from January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.
Is Claude Haiku 4.5 being replaced?
Anthropic announced Claude Haiku 5.5 on September 22, 2026. Haiku 4.5 retires no sooner than October 15, 2026.
How can I see my own cost on each model?
EveryToken prices every request in your local Claude Code and Gemini CLI history at API rates and splits the total by model, cache savings included. The totals are API-equivalent estimates, so a subscription or free tier will charge differently.
Sources
- Anthropic: Pricing
- Anthropic docs: Claude Haiku 4.5
- Anthropic: Introducing Claude Haiku 4.5
- Anthropic: Claude Haiku
- Anthropic docs: Model deprecations
- Cursor docs: Models and pricing
- OpenRouter: Claude Haiku 4.5
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Google: Gemini API pricing
- Google docs: Gemini 3.8 Flash
- Google: Gemini models
- Google: Latest Gemini model
- Google DeepMind: Gemini 3.8 Flash model card
- Google: Gemini 3.8 Flash
- Gemini CLI source: model configuration
- Cursor docs: Gemini 3.8 Flash
- OpenRouter: Gemini 3.8 Flash
- Anthropic: Prompt caching
- Google: Context caching