Skip to content

Model comparison

Claude Haiku 4.5 vs Gemini 3.5 Flash-Lite for sub-agents

Claude Haiku 4.5 and Gemini 3.5 Flash-Lite both target sub-agent work. Haiku 4.5 costs 3.5x as much per cached session, but only 2x as much per output token.

· Prices as of September 28, 2026

  • Claude Haiku 4.5

    Anthropic · Released October 15, 2025

    Anthropic's fastest and cheapest current model, aimed at latency-sensitive work and at coding sub-agents.

    Claude Haiku 4.5 facts and comparisons
  • Gemini 3.5 Flash-Lite

    Google · Released July 21, 2026

    Google's current low-cost, low-latency tier for high-volume and subagent work, recommended with Gemini 3.8 Flash for new projects.

    Gemini 3.5 Flash-Lite facts and comparisons

The short answer

Gemini 3.5 Flash-Lite is the cheaper sub-agent model, at $0.34 against $1.20 on Claude Haiku 4.5 for the example agentic coding session, a 3.5x gap. The gap shrinks to 2x on output-heavy work, because Flash-Lite's $2.50 output rate is half of Haiku 4.5's $5 while its input rate is less than a third. Haiku 4.5 runs in Claude Code, Cursor, and GitHub Copilot, and Flash-Lite runs in Gemini CLI but not in Cursor or Copilot.

Choose Claude Haiku 4.5 if

  • You spawn sub-agents from Claude Code, where Anthropic pitches Haiku 4.5 for multi-agent refactors and migrations.
  • You use Cursor or GitHub Copilot, which offer Haiku 4.5 and not Gemini 3.5 Flash-Lite.
  • You want thinking capped by a token budget you set: Haiku 4.5 uses manual extended thinking.

Choose Gemini 3.5 Flash-Lite if

  • You want the lower cost: $0.30 input and $2.50 output per million against $1 and $5.
  • Your sub-agents need to read more than 200K tokens at once, since Flash-Lite's window is 1.05M.
  • You care about generation speed: Google cites about 350 output tokens per second for Flash-Lite.
  • You use Gemini CLI, where Flash-Lite answers to the flash-lite alias.

Side by side

Specs and prices

FactClaude Haiku 4.5Gemini 3.5 Flash-Lite
MakerAnthropicGoogle
API model idclaude-haiku-4-5gemini-3.5-flash-lite
ReleasedOctober 15, 2025July 21, 2026
StatusCurrentCurrent
Context window200K tokens1.05M tokens
Max output64K tokens65.5K tokens
Open weightsNoNo
Input, per 1M tokens$1$0.30
Cache hit, per 1M$0.10$0.03
Cache write, per 1M$1.25 (5-minute), $2 (1-hour)$0.30 (same as input)
Output, per 1M tokens$5$2.50
Runs inClaude Code, Cursor, OpenCode, OpenRouter, and GitHub CopilotGemini CLI, OpenCode, and OpenRouter

Standard API rates in US dollars, as published by each maker (Claude Haiku 4.5: September 26, 2026; Gemini 3.5 Flash-Lite: September 28, 2026). Batch and priority tiers, taxes, and subscription plans are not included.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadClaude Haiku 4.5Gemini 3.5 Flash-Lite
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$1.20$0.34
Large one-off review, 150K input with no cache hits, 10K output$0.20$0.07
Output-heavy generation, 30K input, 80K output$0.43$0.21
A month of sessions, 110 sessions: 5 a day, 22 working days$132.00$36.85
Where the session’s cost goes
Cache writes$0.65$0.12
Cache reads$0.20$0.06
Uncached input$0.10$0.03
Output$0.25$0.13
caching saves on the session with Claude Haiku 4.5 (56%)
$1.55
caching saves on the session with Gemini 3.5 Flash-Lite (61%)
$0.54

Where the 3.5x session gap comes from

Gemini 3.5 Flash-Lite charges $0.30 per million input tokens, $0.03 per million cache hits, and $2.50 per million output tokens. Claude Haiku 4.5 charges $1 for input, $0.10 for hits, and $5 for output. Input and cache hits are 3.3x apart. Output is only 2x apart.

Cache writes are further apart than any of those. Google publishes no write price, so written tokens cost ordinary input, $0.30 per million. Anthropic charges $1.25 for a 5-minute write and $2 for a 1-hour write. In the example session, which writes 400K tokens and splits them evenly between Claude's two lifetimes, writes cost $0.65 on Haiku 4.5 and $0.12 on Flash-Lite. That $0.53 is most of the session's $0.86 difference.

The totals come to $1.20 and $0.34 per session, and $132.00 against $36.85 over 110 sessions a month. Caching saves $1.55 per session on Haiku 4.5, 56% of the uncached cost, and $0.54 on Flash-Lite, 61%.

Output is a bigger share of Flash-Lite's cost

Flash-Lite's output rate is high relative to its own input rate, $2.50 against $0.30, where Haiku 4.5 pairs $5 output with $1 input. On the example session, output is 38% of Flash-Lite's cost and 21% of Haiku 4.5's. The more a sub-agent writes, the closer the two models get.

The workloads bear that out. The large one-off review, with little output, costs $0.20 on Haiku 4.5 and $0.07 on Flash-Lite, 2.9x. The output-heavy generation costs $0.43 against $0.21, 2x. A sub-agent that mostly reads files and returns short answers sits at the wide end of that range, and one that writes whole files sits at the narrow end.

Thinking settings move output too. Flash-Lite's default thinking level is minimal, and raising it adds output tokens. Haiku 4.5 thinks only when you give it a budget. The tokenizers also differ, so the same file does not produce the same token count on both.

Context, tools, and what comes next for each

Maximum output is nearly the same, 64K on Haiku 4.5 and 65.5K on Flash-Lite. Context is not: 200K against 1.05M, a 5.2x difference in Flash-Lite's favor. For sub-agents that parse long documents, one of the uses Google names, that difference can decide the choice.

Google positions Flash-Lite as its current low-cost, low-latency tier for high-volume and sub-agent work, recommended alongside Gemini 3.8 Flash for new projects, with computer use built in as a tool. Anthropic positions Haiku 4.5 for latency-sensitive work and coding sub-agents, and credits it with coding performance similar to Claude Sonnet 4 for a third of the price.

Haiku 4.5 is older, released in October 2025, and its retirement is listed as not sooner than October 15, 2026. Claude Haiku 5.5 was announced on September 22, 2026. Flash-Lite was released on July 21, 2026, and is a current model.

Prompt caching

How each maker bills cached tokens

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

Google

Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.

Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.

On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.

Source: Google: Context caching

Your own numbers

See what Claude Haiku 4.5 and Gemini 3.5 Flash-Lite really cost you.

everyaitoken reads your Claude Code, Cursor, OpenCode, OpenRouter, and Gemini CLI history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is Gemini 3.5 Flash-Lite cheaper than Claude Haiku 4.5?

Yes, on every rate. The example agentic session costs $0.34 against $1.20, and output-heavy generation $0.21 against $0.43. A month of 110 sessions differs by $95.15.

Which has the larger context window?

Gemini 3.5 Flash-Lite, at 1.05M tokens against 200K for Claude Haiku 4.5. Their maximum outputs are close, 65.5K and 64K.

Can I use Gemini 3.5 Flash-Lite in Cursor or GitHub Copilot?

Neither offers it. It runs in Gemini CLI, OpenRouter, and OpenCode. Claude Haiku 4.5 is offered in Claude Code, Cursor, OpenRouter, OpenCode, and GitHub Copilot.

How do I see what my sub-agents cost?

EveryToken reads your local Claude Code and Gemini CLI history, prices each request at API rates, and shows cost per model with what caching saved. That makes it easy to see how much of a session's cost comes from sub-agents.

  • Claude Haiku 4.5 vs Gemini 3.8 Flash

    Gemini 3.8 Flash undercuts Claude Haiku 4.5 on introductory rates, $0.71 against $1.20 per cached session. From January 1, 2027, its rates top Haiku's.

  • Claude Opus 5.5 vs Claude Haiku 4.5

    Claude Opus 5.5 lists at 4x the rates of Claude Haiku 4.5, yet a cached coding session costs 3.7x. Context and output limits, and Haiku 4.5's retirement date.

  • Claude Sonnet 5 vs Claude Haiku 4.5

    Claude Haiku 4.5 costs half as much as Claude Sonnet 5, with a 200K context window and a retirement date ahead. When the cheaper model fits a coding workflow.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash-Lite

    Google recommends Gemini 3.8 Flash and Gemini 3.5 Flash-Lite together for new projects. Where each fits, the 2.1x session gap, and what changes in 2027.

  • Gemini 3.5 Flash-Lite vs Gemini 3.1 Flash-Lite

    Gemini 3.1 Flash-Lite shuts down on May 7, 2027, and Gemini 3.5 Flash-Lite replaces it at higher rates. What the move costs, mostly on output.

  • Claude Haiku 4.5 vs GPT-5.6 Luna

    After an 80% price cut, GPT-5.6 Luna costs a fifth of Claude Haiku 4.5 for input. A month of cached coding sessions: $24.20 against $132.00.