Model comparison
Claude Haiku 4.5 vs Gemini 3.5 Flash-Lite for sub-agents
Claude Haiku 4.5 and Gemini 3.5 Flash-Lite both target sub-agent work. Haiku 4.5 costs 3.5x as much per cached session, but only 2x as much per output token.
· Prices as of September 28, 2026
Claude Haiku 4.5
Anthropic · Released October 15, 2025
Anthropic's fastest and cheapest current model, aimed at latency-sensitive work and at coding sub-agents.
Claude Haiku 4.5 facts and comparisonsGemini 3.5 Flash-Lite
Google · Released July 21, 2026
Google's current low-cost, low-latency tier for high-volume and subagent work, recommended with Gemini 3.8 Flash for new projects.
Gemini 3.5 Flash-Lite facts and comparisons
The short answer
Gemini 3.5 Flash-Lite is the cheaper sub-agent model, at $0.34 against $1.20 on Claude Haiku 4.5 for the example agentic coding session, a 3.5x gap. The gap shrinks to 2x on output-heavy work, because Flash-Lite's $2.50 output rate is half of Haiku 4.5's $5 while its input rate is less than a third. Haiku 4.5 runs in Claude Code, Cursor, and GitHub Copilot, and Flash-Lite runs in Gemini CLI but not in Cursor or Copilot.
Choose Claude Haiku 4.5 if
- You spawn sub-agents from Claude Code, where Anthropic pitches Haiku 4.5 for multi-agent refactors and migrations.
- You use Cursor or GitHub Copilot, which offer Haiku 4.5 and not Gemini 3.5 Flash-Lite.
- You want thinking capped by a token budget you set: Haiku 4.5 uses manual extended thinking.
Choose Gemini 3.5 Flash-Lite if
- You want the lower cost: $0.30 input and $2.50 output per million against $1 and $5.
- Your sub-agents need to read more than 200K tokens at once, since Flash-Lite's window is 1.05M.
- You care about generation speed: Google cites about 350 output tokens per second for Flash-Lite.
- You use Gemini CLI, where Flash-Lite answers to the flash-lite alias.
Side by side
Specs and prices
| Fact | Claude Haiku 4.5 | Gemini 3.5 Flash-Lite |
|---|---|---|
| Maker | Anthropic | |
| API model id | claude-haiku-4-5 | gemini-3.5-flash-lite |
| Released | October 15, 2025 | July 21, 2026 |
| Status | Current | Current |
| Context window | 200K tokens | 1.05M tokens |
| Max output | 64K tokens | 65.5K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $1 | $0.30 |
| Cache hit, per 1M | $0.10 | $0.03 |
| Cache write, per 1M | $1.25 (5-minute), $2 (1-hour) | $0.30 (same as input) |
| Output, per 1M tokens | $5 | $2.50 |
| Runs in | Claude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot | Gemini CLI, OpenCode, and OpenRouter |
Standard API rates in US dollars, as published by each maker (Claude Haiku 4.5: September 26, 2026; Gemini 3.5 Flash-Lite: September 28, 2026). Batch and priority tiers, taxes, and subscription plans are not included.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Claude Haiku 4.5 | Gemini 3.5 Flash-Lite |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $1.20 | $0.34 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.20 | $0.07 |
| Output-heavy generation, 30K input, 80K output | $0.43 | $0.21 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $132.00 | $36.85 |
| Where the session’s cost goes | ||
| Cache writes | $0.65 | $0.12 |
| Cache reads | $0.20 | $0.06 |
| Uncached input | $0.10 | $0.03 |
| Output | $0.25 | $0.13 |
- caching saves on the session with Claude Haiku 4.5 (56%)
- $1.55
- caching saves on the session with Gemini 3.5 Flash-Lite (61%)
- $0.54
Where the 3.5x session gap comes from
Gemini 3.5 Flash-Lite charges $0.30 per million input tokens, $0.03 per million cache hits, and $2.50 per million output tokens. Claude Haiku 4.5 charges $1 for input, $0.10 for hits, and $5 for output. Input and cache hits are 3.3x apart. Output is only 2x apart.
Cache writes are further apart than any of those. Google publishes no write price, so written tokens cost ordinary input, $0.30 per million. Anthropic charges $1.25 for a 5-minute write and $2 for a 1-hour write. In the example session, which writes 400K tokens and splits them evenly between Claude's two lifetimes, writes cost $0.65 on Haiku 4.5 and $0.12 on Flash-Lite. That $0.53 is most of the session's $0.86 difference.
The totals come to $1.20 and $0.34 per session, and $132.00 against $36.85 over 110 sessions a month. Caching saves $1.55 per session on Haiku 4.5, 56% of the uncached cost, and $0.54 on Flash-Lite, 61%.
Output is a bigger share of Flash-Lite's cost
Flash-Lite's output rate is high relative to its own input rate, $2.50 against $0.30, where Haiku 4.5 pairs $5 output with $1 input. On the example session, output is 38% of Flash-Lite's cost and 21% of Haiku 4.5's. The more a sub-agent writes, the closer the two models get.
The workloads bear that out. The large one-off review, with little output, costs $0.20 on Haiku 4.5 and $0.07 on Flash-Lite, 2.9x. The output-heavy generation costs $0.43 against $0.21, 2x. A sub-agent that mostly reads files and returns short answers sits at the wide end of that range, and one that writes whole files sits at the narrow end.
Thinking settings move output too. Flash-Lite's default thinking level is minimal, and raising it adds output tokens. Haiku 4.5 thinks only when you give it a budget. The tokenizers also differ, so the same file does not produce the same token count on both.
Context, tools, and what comes next for each
Maximum output is nearly the same, 64K on Haiku 4.5 and 65.5K on Flash-Lite. Context is not: 200K against 1.05M, a 5.2x difference in Flash-Lite's favor. For sub-agents that parse long documents, one of the uses Google names, that difference can decide the choice.
Google positions Flash-Lite as its current low-cost, low-latency tier for high-volume and sub-agent work, recommended alongside Gemini 3.8 Flash for new projects, with computer use built in as a tool. Anthropic positions Haiku 4.5 for latency-sensitive work and coding sub-agents, and credits it with coding performance similar to Claude Sonnet 4 for a third of the price.
Haiku 4.5 is older, released in October 2025, and its retirement is listed as not sooner than October 15, 2026. Claude Haiku 5.5 was announced on September 22, 2026. Flash-Lite was released on July 21, 2026, and is a current model.
Prompt caching
How each maker bills cached tokens
Anthropic
Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.
A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.
Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.
In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.
Source: Anthropic: Prompt caching
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what Claude Haiku 4.5 and Gemini 3.5 Flash-Lite really cost you.
everyaitoken reads your Claude Code, Cursor, OpenCode, OpenRouter, and Gemini CLI history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is Gemini 3.5 Flash-Lite cheaper than Claude Haiku 4.5?
Yes, on every rate. The example agentic session costs $0.34 against $1.20, and output-heavy generation $0.21 against $0.43. A month of 110 sessions differs by $95.15.
Which has the larger context window?
Gemini 3.5 Flash-Lite, at 1.05M tokens against 200K for Claude Haiku 4.5. Their maximum outputs are close, 65.5K and 64K.
Can I use Gemini 3.5 Flash-Lite in Cursor or GitHub Copilot?
Neither offers it. It runs in Gemini CLI, OpenRouter, and OpenCode. Claude Haiku 4.5 is offered in Claude Code, Cursor, OpenRouter, OpenCode, and GitHub Copilot.
How do I see what my sub-agents cost?
EveryToken reads your local Claude Code and Gemini CLI history, prices each request at API rates, and shows cost per model with what caching saved. That makes it easy to see how much of a session's cost comes from sub-agents.
Sources
- Anthropic: Pricing
- Anthropic docs: Claude Haiku 4.5
- Anthropic: Introducing Claude Haiku 4.5
- Anthropic: Claude Haiku
- Anthropic docs: Model deprecations
- Cursor docs: Models and pricing
- OpenRouter: Claude Haiku 4.5
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Google: Gemini API pricing
- Google docs: Gemini 3.5 Flash-Lite
- Google: Gemini models
- Google: Gemini 3.6 Flash and 3.5 Flash-Lite
- Gemini CLI source: model configuration
- OpenRouter: Gemini 3.5 Flash-Lite
- Anthropic: Prompt caching
- Google: Context caching