Skip to content

Model comparison

Claude Opus 4.8 to Claude Opus 5.5: the cost of staying

Claude Opus 5.5 lists 20% below Claude Opus 4.8 and bills cache hits at $0.20, not $0.50. What staying costs, and when Anthropic still recommends Opus 4.8.

· Prices as of September 26, 2026

  • Claude Opus 5.5

    Anthropic · Released September 22, 2026

    Anthropic's recommended starting model for most work, built for long-running agentic coding. Anthropic says it matches Claude Fable 5.1 on most work at a much lower price.

    Claude Opus 5.5 facts and comparisons
  • Claude Opus 4.8

    Anthropic · Released May 28, 2026 · Previous generation

    An Opus upgrade over 4.7 focused on judgment and collaboration. Anthropic still recommends it for cybersecurity work that needs reduced guardrails.

    Claude Opus 4.8 facts and comparisons

The short answer

Claude Opus 5.5 costs less than Claude Opus 4.8 on every rate, and the example agentic coding session comes to $4.40 against $6.00, or $176.00 less over a month of 110 sessions. Anthropic recommends Opus 5.5 as its starting model for most work, but still recommends Opus 4.8 for cybersecurity work that needs reduced guardrails.

Choose Claude Opus 5.5 if

  • You're on Opus 4.8 for general coding and want the same Opus tier at 20% lower list prices.
  • You want Claude Code's default model on Pro, Max, Team, and Enterprise plans.
  • You run long sessions with heavy cache reads, where a hit costs $0.20 instead of $0.50.

Choose Claude Opus 4.8 if

  • You do security work that needs reduced guardrails, the case Anthropic still recommends Opus 4.8 for.
  • You have a workflow tuned around Opus 4.8 at xhigh effort and need time to re-validate it.

Side by side

Specs and prices

FactClaude Opus 5.5Claude Opus 4.8
MakerAnthropicAnthropic
API model idclaude-opus-5-5claude-opus-4-8
ReleasedSeptember 22, 2026May 28, 2026
StatusCurrentPrevious generation
Context window1M tokens1M tokens
Max output128K tokens128K tokens
Open weightsNoNo
Input, per 1M tokens$4$5
Cache hit, per 1M$0.20$0.50
Cache write, per 1M$5 (5-minute), $8 (1-hour)$6.25 (5-minute), $10 (1-hour)
Output, per 1M tokens$20$25
Runs inClaude Code, Cursor, OpenCode, OpenRouter, and GitHub CopilotClaude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker on September 26, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Claude Opus 5.5: The full 1M context window is billed at standard rates. Fast mode, a research preview on the Claude API, costs $8 input and $40 output per million tokens. Claude Opus 4.8: The full 1M context window is billed at standard rates. Fast mode, a research preview on the Claude API, costs $10 input and $50 output per million tokens.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadClaude Opus 5.5Claude Opus 4.8
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$4.40$6.00
Large one-off review, 150K input with no cache hits, 10K output$0.80$1.00
Output-heavy generation, 30K input, 80K output$1.72$2.15
A month of sessions, 110 sessions: 5 a day, 22 working days$484.00$660.00
Where the session’s cost goes
Cache writes$2.60$3.25
Cache reads$0.40$1.00
Uncached input$0.40$0.50
Output$1.00$1.25
caching saves on the session with Claude Opus 5.5 (60%)
$6.60
caching saves on the session with Claude Opus 4.8 (56%)
$7.75

What staying on Claude Opus 4.8 costs

Claude Opus 4.8 was released in May 2026 at $5 per million input tokens and $25 per million output tokens. Claude Opus 5.5, released in September, lists at $4 and $20. Cache writes are cheaper by the same 20%: $5 and $8 for 5-minute and 1-hour writes, against $6.25 and $10 on Opus 4.8.

Cache hits fell further. Opus 4.8 bills a hit at 0.1x its input price, $0.50 per million; Opus 5.5 bills 0.05x, $0.20. On the example session, which reads 2M tokens from the cache, that line drops from $1.00 to $0.40, and from 17% of the session to 9%.

Put together, the session costs 27% less on Opus 5.5 while uncached work costs 20% less: the large one-off review is $0.80 against $1.00, and the output-heavy generation $1.72 against $2.15. Over 110 sessions a month that is $484.00 against $660.00. These are API-equivalent estimates. On a Claude subscription you pay the plan price, not per-token rates.

What Anthropic says changed between Opus 4.8 and Opus 5.5

Anthropic measures its Opus 5.5 claims against Claude Opus 5, the release in between, not against Opus 4.8. It says Opus 5.5 outputs more than 30% faster than Opus 5, reads cache 60% cheaper, and performs at the level of Claude Fable 5.1 on most work. Since Opus 5 and Opus 4.8 charge the same $0.50 for a hit, the 60% cache cut applies against Opus 4.8 too.

Anthropic names agentic coding, computer use, and knowledge work as Opus 5.5 strengths, and long, sprawling jobs like codebase-wide migrations and audits in particular. Its pitch for Opus 4.8 was an upgrade over Opus 4.7 focused on judgment and collaboration, with the consistency and autonomy to keep working on long-running tasks. Opus 4.8 is now a legacy model, and the use Anthropic still recommends it for is cybersecurity work that needs reduced guardrails.

Effort, fast mode, and tokens after the switch

Effort defaults differ sharply. Anthropic recommends starting Opus 4.8 at xhigh effort for coding and agentic work, while Opus 5.5 runs adaptive thinking at medium effort by default. Moving a task from xhigh on Opus 4.8 to medium on Opus 5.5 changes the model and the setting at once, so compare cost per finished task rather than per token.

Fast mode, a research preview on the Claude API, costs $10 input and $50 output per million on Opus 4.8, which Anthropic describes as 2.5x speed, and $8 and $40 on Opus 5.5. Both models share Anthropic's newer tokenizer, which counts about 30% more tokens than earlier Claude models for the same text, so switching doesn't change how many tokens a prompt uses. Both take 1M tokens of context at standard rates and write up to 128K.

To see the switch in your own numbers, EveryToken reads your Claude Code history on your Mac, prices each request at API rates, and splits cost and cache savings by model.

Prompt caching

How Anthropic bills cached tokens

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

Your own numbers

See what Claude Opus 5.5 and Claude Opus 4.8 really cost you.

everyaitoken reads your Claude Code, Cursor, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is Claude Opus 5.5 cheaper than Claude Opus 4.8?

Yes. Input, output, and cache writes cost 20% less, and a cache hit costs $0.20 instead of $0.50. The example session is $4.40 on Opus 5.5 and $6.00 on Opus 4.8.

Why would anyone stay on Claude Opus 4.8?

Anthropic still recommends Opus 4.8 for cybersecurity work that needs reduced guardrails. It is a legacy model that stays available.

Does Claude Opus 5.5 use a different tokenizer from Claude Opus 4.8?

No. Both use the newer tokenizer that every Claude model from 4.7 on uses, so the same text counts as the same number of tokens on either.

Which effort setting should I start with on Claude Opus 5.5?

Anthropic sets medium as the default for Opus 5.5 and notes that changing effort keeps the prompt cache. For Opus 4.8, Anthropic recommends starting at xhigh for coding and agentic work.

  • Claude Fable 5.1 vs Claude Opus 5.5

    Claude Fable 5.1 lists at 2.5x the rates of Claude Opus 5.5. What that premium buys, why cheap cache hits barely narrow it, and when Anthropic suggests it.

  • Claude Opus 5 vs Claude Opus 4.8

    Claude Opus 5 and Claude Opus 4.8 cost exactly the same, from $5 input to $0.50 cache hits. How Anthropic positions each, and what it now recommends instead.

  • Claude Opus 5.5 vs Claude Haiku 4.5

    Claude Opus 5.5 lists at 4x the rates of Claude Haiku 4.5, yet a cached coding session costs 3.7x. Context and output limits, and Haiku 4.5's retirement date.

  • Claude Opus 5.5 vs Claude Opus 5

    Claude Opus 5.5 is 20% cheaper per token than Claude Opus 5, and cache hits cost 60% less. What that means for Claude Code sessions, and what stays the same.

  • Claude Opus 5.5 vs Claude Sonnet 5

    Claude Opus 5.5 costs twice as much per token as Claude Sonnet 5, yet cache hits cost $0.20 on both. What that means for agentic coding sessions.

  • Claude Opus 4.8 vs Gemini 3.1 Pro Preview

    Claude Opus 4.8 costs 3x Gemini 3.1 Pro Preview on a cached coding session, $6.00 against $2.00. How long prompts, fast mode, and preview status shift that.