Moonshot AI
Kimi K3: price, context window, and caching
Moonshot's flagship and the largest open model it has released, aimed at long-horizon coding and knowledge work.
Released July 16, 2026 · Prices as of September 28, 2026
In Moonshot AI’s words
“Kimi K3 is Kimi's most capable flagship model to date, with 2.8 trillion parameters.”
What Moonshot AI says it’s good at
Facts
Specs and prices
| Fact | Kimi K3 |
|---|---|
| Maker | Moonshot AI |
| API model id | kimi-k3 |
| Released | July 16, 2026 |
| Status | Current |
| Context window | 1.05M tokens |
| Max output | 1.05M tokens |
| Open weights | Yes |
| Input, per 1M tokens | $3 |
| Cache hit, per 1M | $0.30 |
| Cache write, per 1M | $3 (same as input) |
| Output, per 1M tokens | $15 |
| Runs in | Cursor, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by the maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Kimi K3: A 1-hour cache write costs $6 per million tokens; this table uses the 5-minute default.
Good to know
- Open weights under Moonshot's own Kimi K3 license.
- Output defaults to 131,072 tokens per request and can be raised to the full context window.
- API access unlocks after a minimum $1 top-up.
Cost
What typical work costs
Example token counts at Kimi K3’s published rates. On the agentic session, caching saves $5.40 against billing every token as ordinary input.
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $2.85 |
|---|---|
| Large one-off review, 150K input with no cache hits, 10K output | $0.60 |
| Output-heavy generation, 30K input, 80K output | $1.29 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $313.50 |
Prompt caching
How Moonshot AI bills cached tokens
Moonshot AI
The Kimi API caches prompt prefixes automatically, for 5 minutes by default or 1 hour on request. A cache hit on Kimi K3 costs one-tenth of the input price, and each hit restarts the cache lifetime at no charge.
Kimi bills cache writes separately: $3 per million tokens for the 5-minute cache, the same as ordinary input, and $6 for the 1-hour cache. This blog prices Kimi's writes at the 5-minute default.
Source: Kimi API: Chat pricing
Kimi K3 compared
Grok 4.7 vs Kimi K3
Grok 4.7 and Kimi K3 both run in Cursor and GitHub Copilot. Grok 4.7 costs less on every workload, while Kimi K3 adds open weights and a 1.05M window.
Kimi K3 vs GPT-6 Sol
GPT-6 Sol undercuts Kimi K3 on every rate, and an agentic coding session costs $2.10 against $2.85. Both list 1.05M context, with different limits inside.
Kimi K3 vs Claude Opus 5.5
Kimi K3 lists 25% below Claude Opus 5.5 on input and output, and the gap widens to 35% on a cached coding session. Cache writes, not hits, explain it.
Kimi K3 vs Claude Sonnet 5
Kimi K3 lists 1.5x the rates of Claude Sonnet 5, yet an agentic coding session costs just $0.45 more on it, $2.85 against $2.40. Cache writes explain why.
Kimi K3 vs DeepSeek-V4-Pro
An agentic coding session costs $2.85 on Kimi K3 and $0.95 on DeepSeek-V4-Pro at peak rates, and DeepSeek's off-peak hours cost 50% less again.
Kimi K3 vs GLM-5.3
Kimi K3 and GLM-5.3 are both open-weight flagships. A coding session costs $2.85 on Kimi K3 and $1.44 on GLM-5.3, yet their cache hits nearly match.
Kimi K3 vs GPT-6 Astra
GPT-6 Astra costs 3.3x Kimi K3 per token and 3.7x on an agentic coding session, $10.50 against $2.85. What each maker claims for its top model.
Qwen3.8-Max vs Kimi K3
Qwen3.8-Max undercuts Kimi K3 on every rate, $1.80 against $2.85 for an agentic coding session. Kimi K3 has open weights; the Qwen3.8-Max API model does not.
Your own numbers
See what Kimi K3 really costs you.
everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.