Skip to content

Model comparison

Claude Haiku 4.5 vs GPT-5.6 Luna for low-cost coding work

After an 80% price cut, GPT-5.6 Luna costs a fifth of Claude Haiku 4.5 for input. A month of cached coding sessions: $24.20 against $132.00.

· Prices as of September 28, 2026

  • Claude Haiku 4.5

    Anthropic · Released October 15, 2025

    Anthropic's fastest and cheapest current model, aimed at latency-sensitive work and at coding sub-agents.

    Claude Haiku 4.5 facts and comparisons
  • GPT-5.6 Luna

    OpenAI · Released July 9, 2026 · Previous generation

    The low-cost GPT-5.6 tier for cost-sensitive, high-volume work, which roughly corresponds to earlier GPT-5 nano models.

    GPT-5.6 Luna facts and comparisons

The short answer

GPT-5.6 Luna is the cheaper model: $0.20 input and $1.20 output per million against $1 and $5 for Claude Haiku 4.5, and a month of example agentic coding sessions costs $24.20 against $132.00. Both already have successors in view, since Codex suggests moving from GPT-5.6 Luna to GPT-6 Luna, and Anthropic announced Claude Haiku 5.5 on September 22, 2026. Haiku 4.5 fits Claude Code sub-agents, and GPT-5.6 Luna fits OpenAI pipelines that already call it.

Choose Claude Haiku 4.5 if

  • You run sub-agents in Claude Code, the job Anthropic names for Haiku 4.5 in multi-agent refactors and migrations.
  • You want to budget thinking tokens by hand: Haiku 4.5 uses manual extended thinking with a token budget.
  • You expect to move to Claude Haiku 5.5 when it ships and want your prompts on Anthropic's small model now.

Choose GPT-5.6 Luna if

  • You want the lower rates: input and cache prices are a fifth of Haiku 4.5's, and output is $1.20 against $5.
  • You need more room: a 1.05M context window and 128K of output, against 200K and 64K.
  • Existing code targets gpt-5.6-luna, and you want to leave it alone while it stays in the API.
  • You want every cache write at 1.25x input and a prefix that stays reusable for at least 30 minutes.

Side by side

Specs and prices

FactClaude Haiku 4.5GPT-5.6 Luna
MakerAnthropicOpenAI
API model idclaude-haiku-4-5gpt-5.6-luna
ReleasedOctober 15, 2025July 9, 2026
StatusCurrentPrevious generation
Context window200K tokens1.05M tokens
Max output64K tokens128K tokens
Open weightsNoNo
Input, per 1M tokens$1$0.20
Cache hit, per 1M$0.10$0.02
Cache write, per 1M$1.25 (5-minute), $2 (1-hour)$0.25
Output, per 1M tokens$5$1.20
Runs inClaude Code, Cursor, OpenCode, OpenRouter, and GitHub CopilotCodex, Cursor, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker (Claude Haiku 4.5: September 26, 2026; GPT-5.6 Luna: September 28, 2026). Batch and priority tiers, taxes, and subscription plans are not included. GPT-5.6 Luna: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadClaude Haiku 4.5GPT-5.6 Luna
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$1.20$0.22
Large one-off review, 150K input with no cache hits, 10K output$0.20$0.04
Output-heavy generation, 30K input, 80K output$0.43$0.10
A month of sessions, 110 sessions: 5 a day, 22 working days$132.00$24.20
Where the session’s cost goes
Cache writes$0.65$0.10
Cache reads$0.20$0.04
Uncached input$0.10$0.02
Output$0.25$0.06
caching saves on the session with Claude Haiku 4.5 (56%)
$1.55
caching saves on the session with GPT-5.6 Luna (61%)
$0.34

What the 80% price cut did to this comparison

OpenAI cut GPT-5.6 Luna's price by 80% on July 30, 2026. It now lists $0.20 per million input tokens, $0.02 per million cache hits, $0.25 per million cache writes, and $1.20 per million output tokens. Claude Haiku 4.5 lists $1 for input, $0.10 for hits, $1.25 for a 5-minute write, $2 for a 1-hour write, and $5 for output.

So Haiku 4.5 costs 5x as much on input, cache hits, and standard cache writes, and 4.2x as much on output. The large one-off review costs $0.20 on Haiku 4.5 and $0.04 on Luna. OpenAI places GPT-5.6 Luna roughly where its earlier GPT-5 nano models sat, as the low-cost GPT-5.6 tier for cost-sensitive, high-volume work.

Output narrows the gap, cache writes widen it

Because Luna's output discount is smaller than its input discount, output-heavy work narrows the gap. The output-heavy generation costs $0.43 on Haiku 4.5 against $0.10 on Luna, 4.3x. Output is 27% of Luna's session cost and 21% of Haiku's.

Cache-heavy work widens it. In the example agentic session, half of Haiku 4.5's 400K written tokens go to Anthropic's 1-hour cache at 2x input, while OpenAI bills every write at 1.25x. Writes cost $0.65 on Haiku 4.5 and $0.10 on Luna, and the session lands at $1.20 against $0.22, a 5.5x gap. Over 110 sessions a month that is $132.00 against $24.20, $107.80 apart.

In percentage terms, caching helps Luna slightly more: it saves 61% of the uncached session cost on Luna, $0.34, against 56% on Haiku 4.5, $1.55. In dollars the saving is bigger on Haiku because its rates are higher.

Two small models with successors already named

Neither model is the newest in its line. Codex suggests moving from GPT-5.6 Luna to GPT-6 Luna, and GPT-5.6 Luna stays in the API. On the Anthropic side, Haiku 4.5's retirement date is not sooner than October 15, 2026, and Claude Haiku 5.5 has been announced, as of September 22, 2026. A choice between these two is worth revisiting as the successors take over.

Their limits differ sharply. Haiku 4.5 tops out at 200K tokens of context and 64K of output. GPT-5.6 Luna has 1.05M and 128K, with a long-context rule: requests over 272K input tokens cost 2x for input and cache and 1.5x for output. Haiku 4.5 uses Anthropic's older tokenizer and manual extended thinking, so neither its token counts nor its thinking controls line up with Luna's.

Cursor, OpenRouter, OpenCode, and GitHub Copilot offer both. Haiku 4.5 is also in Claude Code, and GPT-5.6 Luna in Codex.

Prompt caching

How each maker bills cached tokens

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

OpenAI

Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.

From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.

On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.

Source: OpenAI: Prompt caching

Your own numbers

See what Claude Haiku 4.5 and GPT-5.6 Luna really cost you.

everyaitoken reads your Claude Code, Cursor, OpenCode, OpenRouter, and Codex history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

How much cheaper is GPT-5.6 Luna than Claude Haiku 4.5?

About 5x on input, cache hits, and cache writes, and 4.2x on output. The example agentic session costs $0.22 against $1.20, and a month of 110 sessions $24.20 against $132.00.

Should I use GPT-5.6 Luna or GPT-6 Luna?

Codex suggests moving to GPT-6 Luna, which lists lower rates than GPT-5.6 Luna. GPT-5.6 Luna stays in the API for code that already depends on it.

Is Claude Haiku 4.5 being retired?

Not yet. Its retirement is listed as not sooner than October 15, 2026, and until then it remains a current model.

How can I track what my sub-agents cost on each model?

EveryToken reads local history from Claude Code, Codex, Cursor, OpenCode, and OpenRouter, prices every request at API rates, and splits the total by model with cache savings. It is a $9 one-time macOS app.

  • Claude Haiku 4.5 vs GPT-5.6 Terra

    Claude Haiku 4.5 costs about half of GPT-5.6 Terra per token, $1.20 against $2.20 per coding session. Terra offers a 1.05M context window and 128K output.

  • Claude Haiku 4.5 vs GPT-6 Luna

    GPT-6 Luna lists a tenth of Claude Haiku 4.5's rates and a 1.05M context window against 200K. What each costs for sub-agents and cached coding sessions.

  • Claude Opus 5.5 vs Claude Haiku 4.5

    Claude Opus 5.5 lists at 4x the rates of Claude Haiku 4.5, yet a cached coding session costs 3.7x. Context and output limits, and Haiku 4.5's retirement date.

  • Claude Sonnet 5 vs Claude Haiku 4.5

    Claude Haiku 4.5 costs half as much as Claude Sonnet 5, with a 200K context window and a retirement date ahead. When the cheaper model fits a coding workflow.

  • GPT-5.6 Terra vs GPT-5.6 Luna

    GPT-5.6 Terra costs 10x GPT-5.6 Luna on every rate after OpenAI's July 30 price cuts. What each tier is for, and why Codex now sends them different ways.

  • GPT-6 Luna vs GPT-5.6 Luna

    GPT-6 Luna halves the input price of GPT-5.6 Luna, which already had an 80% cut, and trims output further. What the move saves, and what Codex suggests.