Skip to content

Moonshot AI

Kimi K3: price, context window, and caching

Moonshot's flagship and the largest open model it has released, aimed at long-horizon coding and knowledge work.

Released July 16, 2026 · Prices as of September 28, 2026

In Moonshot AI’s words

“Kimi K3 is Kimi's most capable flagship model to date, with 2.8 trillion parameters.”

Kimi API: Kimi K3 quickstart

What Moonshot AI says it’s good at

  • Long engineering tasks with minimal supervision, large codebases, and terminal tools Source
  • Software engineering combined with visual reasoning from screenshots, for frontend and game work Source
  • A cache hit rate above 90% in coding workloads on the official Kimi API Source

Facts

Specs and prices

FactKimi K3
MakerMoonshot AI
API model idkimi-k3
ReleasedJuly 16, 2026
StatusCurrent
Context window1.05M tokens
Max output1.05M tokens
Open weightsYes
Input, per 1M tokens$3
Cache hit, per 1M$0.30
Cache write, per 1M$3 (same as input)
Output, per 1M tokens$15
Runs inCursor, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by the maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Kimi K3: A 1-hour cache write costs $6 per million tokens; this table uses the 5-minute default.

Good to know

  • Open weights under Moonshot's own Kimi K3 license.
  • Output defaults to 131,072 tokens per request and can be raised to the full context window.
  • API access unlocks after a minimum $1 top-up.

Cost

What typical work costs

Example token counts at Kimi K3’s published rates. On the agentic session, caching saves $5.40 against billing every token as ordinary input.

Example workload costs for Kimi K3
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$2.85
Large one-off review, 150K input with no cache hits, 10K output$0.60
Output-heavy generation, 30K input, 80K output$1.29
A month of sessions, 110 sessions: 5 a day, 22 working days$313.50

Prompt caching

How Moonshot AI bills cached tokens

Moonshot AI

The Kimi API caches prompt prefixes automatically, for 5 minutes by default or 1 hour on request. A cache hit on Kimi K3 costs one-tenth of the input price, and each hit restarts the cache lifetime at no charge.

Kimi bills cache writes separately: $3 per million tokens for the 5-minute cache, the same as ordinary input, and $6 for the 1-hour cache. This blog prices Kimi's writes at the 5-minute default.

Source: Kimi API: Chat pricing

Kimi K3 compared

  • Grok 4.7 vs Kimi K3

    Grok 4.7 and Kimi K3 both run in Cursor and GitHub Copilot. Grok 4.7 costs less on every workload, while Kimi K3 adds open weights and a 1.05M window.

  • Kimi K3 vs GPT-6 Sol

    GPT-6 Sol undercuts Kimi K3 on every rate, and an agentic coding session costs $2.10 against $2.85. Both list 1.05M context, with different limits inside.

  • Kimi K3 vs Claude Opus 5.5

    Kimi K3 lists 25% below Claude Opus 5.5 on input and output, and the gap widens to 35% on a cached coding session. Cache writes, not hits, explain it.

  • Kimi K3 vs Claude Sonnet 5

    Kimi K3 lists 1.5x the rates of Claude Sonnet 5, yet an agentic coding session costs just $0.45 more on it, $2.85 against $2.40. Cache writes explain why.

  • Kimi K3 vs DeepSeek-V4-Pro

    An agentic coding session costs $2.85 on Kimi K3 and $0.95 on DeepSeek-V4-Pro at peak rates, and DeepSeek's off-peak hours cost 50% less again.

  • Kimi K3 vs GLM-5.3

    Kimi K3 and GLM-5.3 are both open-weight flagships. A coding session costs $2.85 on Kimi K3 and $1.44 on GLM-5.3, yet their cache hits nearly match.

  • Kimi K3 vs GPT-6 Astra

    GPT-6 Astra costs 3.3x Kimi K3 per token and 3.7x on an agentic coding session, $10.50 against $2.85. What each maker claims for its top model.

  • Qwen3.8-Max vs Kimi K3

    Qwen3.8-Max undercuts Kimi K3 on every rate, $1.80 against $2.85 for an agentic coding session. Kimi K3 has open weights; the Qwen3.8-Max API model does not.

Your own numbers

See what Kimi K3 really costs you.

everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math