Skip to content

Model comparison

Kimi K3 vs Claude Opus 5.5: similar caches, different prices

Kimi K3 lists 25% below Claude Opus 5.5 on input and output, and the gap widens to 35% on a cached coding session. Cache writes, not hits, explain it.

· Prices as of September 28, 2026

  • Kimi K3

    Moonshot AI · Released July 16, 2026

    Moonshot's flagship and the largest open model it has released, aimed at long-horizon coding and knowledge work.

    Kimi K3 facts and comparisons
  • Claude Opus 5.5

    Anthropic · Released September 22, 2026

    Anthropic's recommended starting model for most work, built for long-running agentic coding. Anthropic says it matches Claude Fable 5.1 on most work at a much lower price.

    Claude Opus 5.5 facts and comparisons

The short answer

Kimi K3 costs $2.85 on the example agentic coding session against $4.40 on Claude Opus 5.5, 35% less, even though its cache hits cost $0.30 per million against $0.20. Choose Opus 5.5 for Claude Code, where it is the default, and for Anthropic's long-running agentic work; choose Kimi K3 for lower list prices, open weights, and responses of up to 1.05M tokens.

Choose Kimi K3 if

  • You want list prices 25% lower on both input and output: $3 and $15 per million tokens, against $4 and $20.
  • Your agent writes a lot of fresh context, since a 5-minute cache write on Kimi K3 costs the $3 input rate while Opus 5.5 charges $5.
  • You need very long single outputs: Kimi K3 defaults to 131,072 tokens per request and can be raised to its full 1.05M window.
  • You want open weights, which Moonshot publishes under its own Kimi K3 license.

Choose Claude Opus 5.5 if

  • Claude Code is your daily tool: it is built around Anthropic's models, with Opus 5.5 as the default.
  • Your sessions reread the cache far more than the example does: an Opus 5.5 hit costs $0.20 per million, a third less than Kimi K3's $0.30.
  • You want the option of fast mode on the Claude API, a research preview that raises Opus 5.5 to $8 input and $40 output per million tokens.
  • You run codebase-wide migrations and audits, the long, sprawling jobs Anthropic names as an Opus 5.5 strength.

Side by side

Specs and prices

FactKimi K3Claude Opus 5.5
MakerMoonshot AIAnthropic
API model idkimi-k3claude-opus-5-5
ReleasedJuly 16, 2026September 22, 2026
StatusCurrentCurrent
Context window1.05M tokens1M tokens
Max output1.05M tokens128K tokens
Open weightsYesNo
Input, per 1M tokens$3$4
Cache hit, per 1M$0.30$0.20
Cache write, per 1M$3 (same as input)$5 (5-minute), $8 (1-hour)
Output, per 1M tokens$15$20
Runs inCursor, OpenCode, OpenRouter, and GitHub CopilotClaude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker (Kimi K3: September 28, 2026; Claude Opus 5.5: September 26, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Kimi K3: A 1-hour cache write costs $6 per million tokens; this table uses the 5-minute default. Claude Opus 5.5: The full 1M context window is billed at standard rates. Fast mode, a research preview on the Claude API, costs $8 input and $40 output per million tokens.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadKimi K3Claude Opus 5.5
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$2.85$4.40
Large one-off review, 150K input with no cache hits, 10K output$0.60$0.80
Output-heavy generation, 30K input, 80K output$1.29$1.72
A month of sessions, 110 sessions: 5 a day, 22 working days$313.50$484.00
Where the session’s cost goes
Cache writes$1.20$2.60
Cache reads$0.60$0.40
Uncached input$0.30$0.40
Output$0.75$1.00
caching saves on the session with Kimi K3 (65%)
$5.40
caching saves on the session with Claude Opus 5.5 (60%)
$6.60

Kimi K3 and Claude Opus 5.5 cache alike but bill differently

Moonshot's cache for Kimi K3 works much like Anthropic's for Claude Opus 5.5. Both offer a 5-minute and a 1-hour cache lifetime, and on both a cache hit restarts that lifetime at no charge. Both also bill cache writes separately from ordinary input.

The prices behind those writes differ. Kimi charges $3 per million for a 5-minute write, the same as its input, and $6 for a 1-hour write. Anthropic charges 1.25x input for a 5-minute write and 2x for a 1-hour write, which is $5 and $8 on Opus 5.5. A 1-hour write costs 2x input at both makers, but only Anthropic adds a premium to the 5-minute write.

Hits go the other way. Kimi K3 bills a hit at one-tenth of its input price, $0.30 per million. Anthropic prices Opus 5.5 hits at 0.05x input, $0.20, so they cost 33% less than Kimi's even though every other Opus 5.5 rate is higher.

Why the gap widens from 25% to 35% on a session

On uncached work the list-price ratio holds. The large one-off review costs $0.60 on Kimi K3 and $0.80 on Opus 5.5, and the output-heavy generation $1.29 against $1.72, both 25% less on Kimi K3.

In the agentic session, Kimi K3 and Opus 5.5 each write 400K tokens to the cache and read 2M. Writes cost $1.20 on Kimi K3 and $2.60 on Opus 5.5, a $1.40 difference that makes up most of the $1.55 total gap. Reads claw back $0.20 for Opus 5.5, at $0.40 against $0.60. The session ends at $2.85 against $4.40, or $313.50 against $484.00 over 110 sessions a month.

One assumption shapes that result. This comparison prices all of Kimi's writes at the 5-minute default and splits Anthropic's evenly between the 5-minute and 1-hour lifetimes. A Kimi setup that asks for the 1-hour cache pays $6 per million for those writes, which narrows the gap. Claude Code picks the Opus 5.5 lifetime for you: 1-hour writes on a Claude subscription, 5-minute writes with an API key. The cache example on our homepage walks through the session.

Output limits, open weights, and where each model runs

Kimi K3's output limit is its whole 1.05M context window. Kimi K3 requests default to 131,072 output tokens and can be raised from there. Opus 5.5 writes up to 128K per request. Both take about 1M tokens of context.

Cursor, OpenCode, OpenRouter, and GitHub Copilot offer both models, so in those tools moving between Kimi K3 and Opus 5.5 is a menu choice. Claude Code, Anthropic's own agent, is built around Anthropic's models and defaults to Opus 5.5. Kimi K3 is also sold on Moonshot's own API, which unlocks after a minimum $1 top-up.

Kimi K3 comes with open weights under Moonshot's own Kimi K3 license, so self-hosting is possible, though this page does not estimate what that would cost. The Kimi prices here are Moonshot's own API rates. Through OpenRouter, Kimi K3 requests go to one of several providers, whose prices can differ.

What Moonshot and Anthropic say about their models

Moonshot says "Kimi K3 is Kimi's most capable flagship model to date, with 2.8 trillion parameters," and it is the largest open model Moonshot has released. It points to long engineering tasks with minimal supervision, large codebases, and terminal tools, and to software engineering combined with visual reasoning from screenshots for frontend and game work. Moonshot adds that Kimi K3 sees a cache hit rate above 90% in coding workloads on the official Kimi API.

Anthropic positions Opus 5.5 as its recommended starting model for most work and says it performs at the level of Claude Fable 5.1 on most work. It singles out long, sprawling jobs such as codebase-wide migrations and audits, and says Opus 5.5 output is more than 30% faster than Claude Opus 5. Adaptive thinking is always on, at medium effort by default.

EveryToken prices Opus 5.5 at Anthropic's rates in Claude Code, Cursor, and OpenCode. It prices Kimi K3 when you use it through OpenRouter, from OpenRouter's catalog, and not when you call Moonshot's API directly.

Prompt caching

How each maker bills cached tokens

Moonshot AI

The Kimi API caches prompt prefixes automatically, for 5 minutes by default or 1 hour on request. A cache hit on Kimi K3 costs one-tenth of the input price, and each hit restarts the cache lifetime at no charge.

Kimi bills cache writes separately: $3 per million tokens for the 5-minute cache, the same as ordinary input, and $6 for the 1-hour cache. This blog prices Kimi's writes at the 5-minute default.

Source: Kimi API: Chat pricing

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

Your own numbers

See what Kimi K3 and Claude Opus 5.5 really cost you.

everyaitoken reads your OpenRouter, Claude Code, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is Kimi K3 cheaper than Claude Opus 5.5?

Yes. Its input and output rates are 25% lower, and the example agentic session costs $2.85 against $4.40. Opus 5.5 is cheaper on one rate, the cache hit, at $0.20 against $0.30 per million.

Does Kimi K3 charge for cache writes?

Yes, Kimi bills writes separately: $3 per million for the 5-minute cache, the same as its input price, and $6 for the 1-hour cache. Opus 5.5 charges $5 and $8 for the same two lifetimes.

Can I use Kimi K3 and Claude Opus 5.5 in the same tool?

Yes. Cursor, OpenCode, OpenRouter, and GitHub Copilot list both. Claude Code, Anthropic's own agent, uses Opus 5.5 as its default.

How long can a Kimi K3 response be?

Up to its full 1.05M context window. The default is 131,072 output tokens per request, which you can raise. Opus 5.5 writes up to 128K tokens per request.

  • Claude Fable 5.1 vs Claude Opus 5.5

    Claude Fable 5.1 lists at 2.5x the rates of Claude Opus 5.5. What that premium buys, why cheap cache hits barely narrow it, and when Anthropic suggests it.

  • Claude Opus 5.5 vs Claude Opus 4.8

    Claude Opus 5.5 lists 20% below Claude Opus 4.8 and bills cache hits at $0.20, not $0.50. What staying costs, and when Anthropic still recommends Opus 4.8.

  • Claude Opus 5.5 vs Claude Haiku 4.5

    Claude Opus 5.5 lists at 4x the rates of Claude Haiku 4.5, yet a cached coding session costs 3.7x. Context and output limits, and Haiku 4.5's retirement date.

  • Claude Opus 5.5 vs Claude Opus 5

    Claude Opus 5.5 is 20% cheaper per token than Claude Opus 5, and cache hits cost 60% less. What that means for Claude Code sessions, and what stays the same.

  • Claude Opus 5.5 vs Claude Sonnet 5

    Claude Opus 5.5 costs twice as much per token as Claude Sonnet 5, yet cache hits cost $0.20 on both. What that means for agentic coding sessions.

  • Kimi K3 vs Claude Sonnet 5

    Kimi K3 lists 1.5x the rates of Claude Sonnet 5, yet an agentic coding session costs just $0.45 more on it, $2.85 against $2.40. Cache writes explain why.