Skip to content

Anthropic

Claude Haiku 4.5: price, context window, and caching

Anthropic's fastest and cheapest current model, aimed at latency-sensitive work and at coding sub-agents.

Released October 15, 2025 · Prices as of September 26, 2026

In Anthropic’s words

“The fastest model with near-frontier intelligence”

Anthropic docs: Claude Haiku 4.5

What Anthropic says it’s good at

  • Coding performance similar to Claude Sonnet 4 at one-third the cost and more than twice the speed Source
  • Powering sub-agents in multi-agent refactors and migrations Source

Facts

Specs and prices

FactClaude Haiku 4.5
MakerAnthropic
API model idclaude-haiku-4-5
ReleasedOctober 15, 2025
StatusCurrent
Context window200K tokens
Max output64K tokens
Open weightsNo
Input, per 1M tokens$1
Cache hit, per 1M$0.10
Cache write, per 1M$1.25 (5-minute), $2 (1-hour)
Output, per 1M tokens$5
Runs inClaude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by the maker on September 26, 2026. Batch and priority tiers, taxes, and subscription plans are not included.

Good to know

  • It uses Anthropic's older tokenizer, so the same text counts as fewer tokens than on Claude models from 4.7 on, which produce about 30% more.
  • Anthropic lists its retirement as not sooner than October 15, 2026. On September 22, 2026, it said Claude Haiku 5.5 would follow in the coming weeks.
  • It uses manual extended thinking with a token budget, not adaptive thinking or effort levels.

Cost

What typical work costs

Example token counts at Claude Haiku 4.5’s published rates. On the agentic session, caching saves $1.55 against billing every token as ordinary input.

Example workload costs for Claude Haiku 4.5
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$1.20
Large one-off review, 150K input with no cache hits, 10K output$0.20
Output-heavy generation, 30K input, 80K output$0.43
A month of sessions, 110 sessions: 5 a day, 22 working days$132.00

Prompt caching

How Anthropic bills cached tokens

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

Claude Haiku 4.5 compared

  • Claude Haiku 4.5 vs Gemini 3.5 Flash-Lite

    Claude Haiku 4.5 and Gemini 3.5 Flash-Lite both target sub-agent work. Haiku 4.5 costs 3.5x as much per cached session, but only 2x as much per output token.

  • Claude Haiku 4.5 vs Gemini 3.8 Flash

    Gemini 3.8 Flash undercuts Claude Haiku 4.5 on introductory rates, $0.71 against $1.20 per cached session. From January 1, 2027, its rates top Haiku's.

  • Claude Haiku 4.5 vs GPT-5.6 Luna

    After an 80% price cut, GPT-5.6 Luna costs a fifth of Claude Haiku 4.5 for input. A month of cached coding sessions: $24.20 against $132.00.

  • Claude Haiku 4.5 vs GPT-5.6 Terra

    Claude Haiku 4.5 costs about half of GPT-5.6 Terra per token, $1.20 against $2.20 per coding session. Terra offers a 1.05M context window and 128K output.

  • Claude Haiku 4.5 vs GPT-6 Luna

    GPT-6 Luna lists a tenth of Claude Haiku 4.5's rates and a 1.05M context window against 200K. What each costs for sub-agents and cached coding sessions.

  • Claude Opus 5.5 vs Claude Haiku 4.5

    Claude Opus 5.5 lists at 4x the rates of Claude Haiku 4.5, yet a cached coding session costs 3.7x. Context and output limits, and Haiku 4.5's retirement date.

  • Claude Sonnet 5 vs Claude Haiku 4.5

    Claude Haiku 4.5 costs half as much as Claude Sonnet 5, with a 200K context window and a retirement date ahead. When the cheaper model fits a coding workflow.

  • DeepSeek-V4.1-Flash vs Claude Haiku 4.5

    DeepSeek-V4.1-Flash holds 5x the context of Claude Haiku 4.5 and costs $0.22 against $1.20 for a cached coding session. Haiku 4.5's successor is announced.

  • GLM-5.3-Flash vs Claude Haiku 4.5

    GLM-5.3-Flash charges $0.50 per million output tokens against $5 on Claude Haiku 4.5, and $0.16 against $1.20 for a cached coding session. Where each fits.

Your own numbers

See what Claude Haiku 4.5 really costs you.

everyaitoken reads your Claude Code, Cursor, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math