Skip to content

Model comparison

Kimi K3 vs Claude Sonnet 5: when the open model costs more

Kimi K3 lists 1.5x the rates of Claude Sonnet 5, yet an agentic coding session costs just $0.45 more on it, $2.85 against $2.40. Cache writes explain why.

· Prices as of September 28, 2026

  • Kimi K3

    Moonshot AI · Released July 16, 2026

    Moonshot's flagship and the largest open model it has released, aimed at long-horizon coding and knowledge work.

    Kimi K3 facts and comparisons
  • Claude Sonnet 5

    Anthropic · Released June 30, 2026

    Anthropic's balance of speed and intelligence, and a drop-in upgrade for Claude Sonnet 4.6. Anthropic pitches it as close to Claude Opus 4.8 at a lower price.

    Claude Sonnet 5 facts and comparisons

The short answer

Claude Sonnet 5 is the cheaper model here: Kimi K3's input, output, and cache-hit rates are all 1.5x higher, and the example agentic coding session costs $2.85 on Kimi K3 against $2.40 on Sonnet 5. The session gap is only 19% because Kimi's 5-minute cache writes cost less than Anthropic's mix of 5-minute and 1-hour writes, so pick Kimi K3 for open weights, very long outputs, or Moonshot's flagship, and Sonnet 5 for price and Claude Code.

Choose Kimi K3 if

  • You want Moonshot's flagship, which it describes as its most capable model to date, with 2.8 trillion parameters.
  • You need outputs longer than 128K, since Kimi K3 can raise its per-request output to its full 1.05M context window.
  • You want open weights under Moonshot's Kimi K3 license, for self-hosting or for a choice of OpenRouter providers.

Choose Claude Sonnet 5 if

  • You want the lower price on every rate: $2 input and $10 output per million tokens, against $3 and $15.
  • You code in Claude Code, which is built around Anthropic's models and maps its sonnet alias to Sonnet 5.
  • Most of your requests are uncached, where the full 1.5x gap applies: the large review costs $0.40 against $0.60.
  • You value what Anthropic stresses for Sonnet 5, planning and driving browsers and terminals on its own, over Moonshot's screenshot-to-frontend pitch for Kimi K3.

Side by side

Specs and prices

FactKimi K3Claude Sonnet 5
MakerMoonshot AIAnthropic
API model idkimi-k3claude-sonnet-5
ReleasedJuly 16, 2026June 30, 2026
StatusCurrentCurrent
Context window1.05M tokens1M tokens
Max output1.05M tokens128K tokens
Open weightsYesNo
Input, per 1M tokens$3$2
Cache hit, per 1M$0.30$0.20
Cache write, per 1M$3 (same as input)$2.50 (5-minute), $4 (1-hour)
Output, per 1M tokens$15$10
Runs inCursor, OpenCode, OpenRouter, and GitHub CopilotClaude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker (Kimi K3: September 28, 2026; Claude Sonnet 5: September 26, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Kimi K3: A 1-hour cache write costs $6 per million tokens; this table uses the 5-minute default. Claude Sonnet 5: The full 1M context window is billed at standard rates. The launch price became the standard price on August 10, 2026, and a planned increase was cancelled.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadKimi K3Claude Sonnet 5
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$2.85$2.40
Large one-off review, 150K input with no cache hits, 10K output$0.60$0.40
Output-heavy generation, 30K input, 80K output$1.29$0.86
A month of sessions, 110 sessions: 5 a day, 22 working days$313.50$264.00
Where the session’s cost goes
Cache writes$1.20$1.30
Cache reads$0.60$0.40
Uncached input$0.30$0.20
Output$0.75$0.50
caching saves on the session with Kimi K3 (65%)
$5.40
caching saves on the session with Claude Sonnet 5 (56%)
$3.10

Why a 1.5x price gap shrinks to 1.2x on a session

Kimi K3 charges $3 per million input tokens, $0.30 per million cache hits, and $15 per million output tokens. Claude Sonnet 5 charges $2, $0.20, and $10. Each is 50% higher on Kimi K3, so uncached work shows the full 1.5x gap: $0.60 against $0.40 for the large one-off review, and $1.29 against $0.86 for the output-heavy generation.

Cache writes break the pattern. Kimi bills a 5-minute write at $3 per million, the same as its input. Anthropic bills a 5-minute write at 1.25x input, $2.50 on Sonnet 5, and a 1-hour write at 2x, which is $4. The example session writes 400K tokens, split evenly between Anthropic's two lifetimes and priced entirely at Kimi's 5-minute default, so writes cost $1.20 on Kimi K3 and $1.30 on Sonnet 5.

That $0.10 in Kimi's favor offsets part of the other lines: $0.20 more for cache reads, $0.25 more for output, and $0.10 more for fresh input on Kimi K3. The session totals $2.85 against $2.40, a $0.45 gap, or $313.50 against $264.00 over 110 sessions a month. Caching saves 65% on Kimi K3 and 56% on Sonnet 5, because more of Sonnet's cached spend goes to write premiums.

How your cache setup moves the gap

The even split between Anthropic's write lifetimes is this page's assumption, and real Sonnet 5 setups vary. Which lifetime Sonnet 5 writes at depends on how Claude Code is signed in: a Claude subscription uses the 1-hour cache, an API key the 5-minute one. With only 5-minute writes, Sonnet 5 pays $2.50 per million to write the cache, below Kimi's $3, and the session gap moves back toward 1.5x.

On Kimi's side, caching is automatic with a 5-minute default lifetime and a 1-hour option, and a hit resets the clock for free. The 1-hour cache costs $6 per million to write, 2x Kimi's input, the same multiple Anthropic uses. Moonshot reports a cache hit rate above 90% in coding workloads on its official API, and on either model a high hit rate matters, because cache reads make up most of a long session's tokens.

The cache section on our homepage lays out the session behind these Kimi K3 and Sonnet 5 figures. Every number here is an API-equivalent estimate at each maker's own API rate. OpenRouter sends Kimi K3 requests to one of several providers, whose prices can differ from Moonshot's.

A flagship against a mid-tier model

The two sit at different points in their makers' lineups. Moonshot calls Kimi K3 "Kimi's most capable flagship model to date, with 2.8 trillion parameters," and it is the largest open model Moonshot has released. Moonshot pitches it at long engineering tasks with minimal supervision, large codebases, and terminal tools, and at pairing code with visual reasoning from screenshots for frontend and game work.

Anthropic calls Sonnet 5 its balance of speed and intelligence and says it is "built to be the most agentic Sonnet model yet." It claims performance close to Claude Opus 4.8 at lower prices and describes Sonnet 5 as a drop-in upgrade for Claude Sonnet 4.6. Adaptive thinking is on by default, at high effort.

Neither pitch tells you how many tokens Kimi K3 or Sonnet 5 will spend on your tasks. Tokenizers differ, and Anthropic says Sonnet 5 counts about 1.0 to 1.35x as many tokens as Claude Sonnet 4.6 for the same text. Try both on real work before moving a team.

Tools, output limits, and open weights

Cursor, OpenCode, OpenRouter, and GitHub Copilot offer both models. Sonnet 5 is also in Claude Code, which is built around Anthropic's models. Kimi K3 is sold on Moonshot's own API as well, which unlocks after a minimum $1 top-up.

Kimi K3 can write far more in one response. Its output defaults to 131,072 tokens per request and can be raised to the full 1.05M context window, while Sonnet 5 stops at 128K. Context is similar, 1.05M against 1M. Kimi K3 also ships open weights under Moonshot's own license, and this page does not estimate what self-hosting would cost.

EveryToken prices Sonnet 5 at Anthropic's rates in Claude Code, Cursor, and OpenCode, and prices Kimi K3 from OpenRouter's catalog when you use it through OpenRouter.

Prompt caching

How each maker bills cached tokens

Moonshot AI

The Kimi API caches prompt prefixes automatically, for 5 minutes by default or 1 hour on request. A cache hit on Kimi K3 costs one-tenth of the input price, and each hit restarts the cache lifetime at no charge.

Kimi bills cache writes separately: $3 per million tokens for the 5-minute cache, the same as ordinary input, and $6 for the 1-hour cache. This blog prices Kimi's writes at the 5-minute default.

Source: Kimi API: Chat pricing

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

Your own numbers

See what Kimi K3 and Claude Sonnet 5 really cost you.

everyaitoken reads your OpenRouter, Claude Code, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is Kimi K3 more expensive than Claude Sonnet 5?

Yes. Kimi K3's input, output, and cache-hit rates are 1.5x Sonnet 5's, and the example agentic session costs $2.85 against $2.40. The session gap is smaller than the list gap because Kimi's 5-minute cache writes cost less than Anthropic's 1-hour writes.

Which model is cheaper for cache writes?

It depends on the lifetime. For a 5-minute write, Sonnet 5 charges $2.50 per million and Kimi K3 $3. For a 1-hour write, Sonnet 5 charges $4 and Kimi K3 $6.

Which tools list both Kimi K3 and Claude Sonnet 5?

Cursor, OpenCode, OpenRouter, and GitHub Copilot list both. Claude Code is built around Anthropic's models and resolves its sonnet alias to Sonnet 5.

Which model writes longer responses?

Kimi K3. It can raise its per-request output to its full 1.05M context window, about 8.2x the 128K limit on Sonnet 5. Longer outputs cost more, though, at $15 per million output tokens.

  • Claude Fable 5.1 vs Claude Sonnet 5

    Claude Fable 5.1 lists at 5x the rates of Claude Sonnet 5, but cache hits cost $0.25 against $0.20. Why sessions cost 4.4x, and who Anthropic aims Fable 5.1 at.

  • Claude Opus 5.5 vs Claude Sonnet 5

    Claude Opus 5.5 costs twice as much per token as Claude Sonnet 5, yet cache hits cost $0.20 on both. What that means for agentic coding sessions.

  • Claude Sonnet 5 vs Claude Haiku 4.5

    Claude Haiku 4.5 costs half as much as Claude Sonnet 5, with a 200K context window and a retirement date ahead. When the cheaper model fits a coding workflow.

  • Claude Sonnet 5 vs Claude Sonnet 4.6

    Claude Sonnet 5 costs 33% less per token than Claude Sonnet 4.6, but its tokenizer counts up to 1.35x as many tokens. What the upgrade saves in practice.

  • Kimi K3 vs Claude Opus 5.5

    Kimi K3 lists 25% below Claude Opus 5.5 on input and output, and the gap widens to 35% on a cached coding session. Cache writes, not hits, explain it.

  • GLM-5.3-Flash vs Claude Sonnet 5

    GLM-5.3-Flash costs $0.16 for a cached coding session that costs $2.40 on Claude Sonnet 5. What drives a 15x gap, and what price alone can't tell you.