Model comparison
Claude Haiku 4.5 vs GPT-6 Luna: a 10x gap in token prices
GPT-6 Luna lists a tenth of Claude Haiku 4.5's rates and a 1.05M context window against 200K. What each costs for sub-agents and cached coding sessions.
· Prices as of September 28, 2026
Claude Haiku 4.5
Anthropic · Released October 15, 2025
Anthropic's fastest and cheapest current model, aimed at latency-sensitive work and at coding sub-agents.
Claude Haiku 4.5 facts and comparisonsGPT-6 Luna
OpenAI · Released September 22, 2026
The cheapest GPT-6 model, pitched for focused, high-volume, repeatable work, including narrower coding tasks.
GPT-6 Luna facts and comparisons
The short answer
GPT-6 Luna costs a tenth of Claude Haiku 4.5 per token, and a month of example agentic coding sessions comes to $11.55 against $132.00. Luna also offers a 1.05M context window and 128K of output against Haiku 4.5's 200K and 64K. Haiku 4.5 still earns its place as a sub-agent model inside Claude Code, though Anthropic lists its retirement as not sooner than October 15, 2026, with Claude Haiku 5.5 announced.
Choose Claude Haiku 4.5 if
- You run sub-agents inside Claude Code, and Anthropic pitches Haiku 4.5 for multi-agent refactors and migrations.
- You prefer to cap thinking with an explicit token budget: Haiku 4.5 uses manual extended thinking rather than effort levels.
Choose GPT-6 Luna if
- Cost is the constraint: every GPT-6 Luna rate is a tenth of Haiku 4.5's, from $0.10 against $1 for input to $0.50 against $5 for output.
- A job outgrows Claude Haiku 4.5's limits of 200K tokens in and 64K tokens out.
- You work in Codex, whose docs recommend Luna for focused, repeatable tasks, and Free and Go plans get it in the Codex app.
- You want the newer model, released in September 2026, rather than one whose retirement window is already announced.
Side by side
Specs and prices
| Fact | Claude Haiku 4.5 | GPT-6 Luna |
|---|---|---|
| Maker | Anthropic | OpenAI |
| API model id | claude-haiku-4-5 | gpt-6-luna |
| Released | October 15, 2025 | September 22, 2026 |
| Status | Current | Current |
| Context window | 200K tokens | 1.05M tokens |
| Max output | 64K tokens | 128K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $1 | $0.10 |
| Cache hit, per 1M | $0.10 | $0.01 |
| Cache write, per 1M | $1.25 (5-minute), $2 (1-hour) | $0.125 |
| Output, per 1M tokens | $5 | $0.50 |
| Runs in | Claude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot | Codex, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker (Claude Haiku 4.5: September 26, 2026; GPT-6 Luna: September 28, 2026). Batch and priority tiers, taxes, and subscription plans are not included. GPT-6 Luna: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Claude Haiku 4.5 | GPT-6 Luna |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $1.20 | $0.11 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.20 | $0.02 |
| Output-heavy generation, 30K input, 80K output | $0.43 | $0.04 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $132.00 | $11.55 |
| Where the session’s cost goes | ||
| Cache writes | $0.65 | $0.05 |
| Cache reads | $0.20 | $0.02 |
| Uncached input | $0.10 | $0.01 |
| Output | $0.25 | $0.03 |
- caching saves on the session with Claude Haiku 4.5 (56%)
- $1.55
- caching saves on the session with GPT-6 Luna (61%)
- $0.17
A tenth of the price on every token type
GPT-6 Luna charges $0.10 per million input tokens, $0.01 per million cache hits, $0.125 per million cache writes, and $0.50 per million output tokens. Claude Haiku 4.5 charges $1 for input, $0.10 for hits, $1.25 for 5-minute writes, and $5 for output. Every one of those is exactly 10x, which makes Luna 90% cheaper on the rate card.
At these prices, single jobs round to a few cents. The large one-off review costs $0.20 on Haiku 4.5 and $0.02 on Luna, and the output-heavy generation $0.43 against $0.04. The example agentic session costs $1.20 against $0.11, a $1.09 difference.
Why the monthly gap reads 11.4x
A month of 110 sessions costs $132.00 on Haiku 4.5 and $11.55 on Luna, which is 11.4x rather than 10x. Two things explain the extra. Anthropic's 1-hour cache write costs 2x input, $2 per million on Haiku 4.5, while OpenAI charges 1.25x input for every write. The example session puts half of its 400K written tokens in the 1-hour cache on Claude, so writes cost $0.65 on Haiku 4.5 against $0.05 on Luna.
The single-session figure of $0.11 is also rounded up to the cent, which is why one session and one month give slightly different ratios. At prices this low, the monthly figure is the more useful one.
Caching still pays on both. It saves $1.55 per session on Haiku 4.5, 56% of the uncached cost, and $0.17 on Luna, 61%. Anthropic's rule of thumb is that a 5-minute write pays for itself after one read and a 1-hour write after two.
Room to work: 200K against 1.05M tokens
The limits differ more than the names suggest. Haiku 4.5 has a 200K context window and writes up to 64K tokens. Luna has 1.05M and writes up to 128K. A sub-agent that must read a large module, a long log, or several files at once may fit in Luna's window and not in Haiku's.
Luna's price notes add one rule: any request over 272K input tokens costs 2x for input and cache and 1.5x for output, for the whole request. Haiku 4.5 cannot take a prompt that size at all, so the rule only matters when you use Luna's extra room.
Token counts are not directly comparable either. Haiku 4.5 uses Anthropic's older tokenizer, and Claude models from 4.7 on count about 30% more tokens for the same text. OpenAI's tokenizer is different again.
Retirement dates and where each runs
Haiku 4.5 dates from October 2025, and Anthropic has set its retirement for no sooner than October 15, 2026. Anthropic announced Claude Haiku 5.5 on September 22, 2026. Anthropic positions Haiku 4.5 for latency-sensitive work and coding sub-agents, and says its coding performance is similar to Claude Sonnet 4 at one-third the cost and more than twice the speed. It runs in Claude Code, Cursor, OpenRouter, OpenCode, and GitHub Copilot.
GPT-6 Luna was released on September 22, 2026, as the low-cost model of the GPT-6 family. OpenAI calls it "our most efficient model for focused, high-volume tasks." Its reasoning effort defaults to medium and goes up to max in Codex. It runs in Codex, though not in Codex cloud, and on OpenRouter, OpenCode, and GitHub Copilot.
Prompt caching
How each maker bills cached tokens
Anthropic
Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.
A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.
Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.
In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.
Source: Anthropic: Prompt caching
OpenAI
Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.
From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.
On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.
Source: OpenAI: Prompt caching
Your own numbers
See what Claude Haiku 4.5 and GPT-6 Luna really cost you.
everyaitoken reads your Claude Code, Cursor, OpenCode, OpenRouter, and Codex history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Does Claude Haiku 4.5 really cost ten times as much as GPT-6 Luna?
Per token, yes: each GPT-6 Luna rate for input, output, cache hits, and cache writes is a tenth of Haiku 4.5's. On a month of example agentic sessions the gap is 11.4x, $11.55 against $132.00, because Anthropic's 1-hour cache writes cost 2x input.
When will Claude Haiku 4.5 be retired?
Anthropic lists its retirement as not sooner than October 15, 2026. It announced Claude Haiku 5.5 on September 22, 2026.
Can GPT-6 Luna read more code at once than Claude Haiku 4.5?
Yes. Luna's context window is 1.05M tokens against Haiku 4.5's 200K, and it writes up to 128K output tokens against 64K. Requests over 272K input tokens on Luna are billed at long-context rates.
How can I compare what the two cost me?
EveryToken reads your local Claude Code and Codex history, prices each request at API rates, and shows cost and cache savings per model. Its figures are estimates at published rates, not what a subscription charges.
Sources
- Anthropic: Pricing
- Anthropic docs: Claude Haiku 4.5
- Anthropic: Introducing Claude Haiku 4.5
- Anthropic: Claude Haiku
- Anthropic docs: Model deprecations
- Cursor docs: Models and pricing
- OpenRouter: Claude Haiku 4.5
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- OpenAI: API pricing
- OpenAI docs: GPT-6 Luna
- OpenAI: Introducing GPT-6 Sol and Luna
- OpenAI: API changelog
- Codex docs: Models
- OpenRouter: GPT-6 Luna
- Anthropic: Prompt caching
- OpenAI: Prompt caching