Model comparison
Qwen3.8-Max vs Kimi K3: a closed API against open weights
Qwen3.8-Max undercuts Kimi K3 on every rate, $1.80 against $2.85 for an agentic coding session. Kimi K3 has open weights; the Qwen3.8-Max API model does not.
· Prices as of September 28, 2026
Qwen3.8-Max
Alibaba Qwen · Released August 2, 2026
Qwen's most capable model, built for long autonomous coding and professional work.
Qwen3.8-Max facts and comparisonsKimi K3
Moonshot AI · Released July 16, 2026
Moonshot's flagship and the largest open model it has released, aimed at long-horizon coding and knowledge work.
Kimi K3 facts and comparisons
The short answer
Qwen3.8-Max is cheaper on every rate, and the example agentic coding session costs $1.80 on it against $2.85 on Kimi K3, 37% less. Pick Kimi K3 if you want open weights, access in Cursor or GitHub Copilot, or responses longer than 131K tokens; pick Qwen3.8-Max if a hosted API is all you need and you want the lower price through OpenRouter, OpenCode, or Qwen Cloud.
Choose Qwen3.8-Max if
- You want the lower price everywhere: $2 input, $0.25 cache hits, and $6 output per million, against $3, $0.30, and $15.
- Your work is output-heavy, where the gap is widest: the example generation costs $0.54 on Qwen3.8-Max against $1.29.
- Your tool can mark cache breakpoints, where Qwen's explicit cache reads at $0.17 per million, well under Kimi K3's $0.30.
Choose Kimi K3 if
- You need open weights for the exact model you run: Moonshot publishes Kimi K3's, while Qwen keeps the Max API model closed.
- You choose models in Cursor or GitHub Copilot, which offer Kimi K3 and do not offer Qwen3.8-Max.
- You need responses longer than 131K tokens, since Kimi K3 can raise output to its full 1.05M window.
Side by side
Specs and prices
| Fact | Qwen3.8-Max | Kimi K3 |
|---|---|---|
| Maker | Alibaba Qwen | Moonshot AI |
| API model id | qwen3.8-max | kimi-k3 |
| Released | August 2, 2026 | July 16, 2026 |
| Status | Current | Current |
| Context window | 1M tokens | 1.05M tokens |
| Max output | 131K tokens | 1.05M tokens |
| Open weights | No | Yes |
| Input, per 1M tokens | $2 | $3 |
| Cache hit, per 1M | $0.25 | $0.30 |
| Cache write, per 1M | $2 (same as input) | $3 (same as input) |
| Output, per 1M tokens | $6 | $15 |
| Runs in | OpenCode and OpenRouter | Cursor, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Qwen3.8-Max: Qwen Cloud prices. Alibaba Cloud Model Studio's international scope lists the same input and output rates. Kimi K3: A 1-hour cache write costs $6 per million tokens; this table uses the 5-minute default.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Qwen3.8-Max | Kimi K3 |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $1.80 | $2.85 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.36 | $0.60 |
| Output-heavy generation, 30K input, 80K output | $0.54 | $1.29 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $198.00 | $313.50 |
| Where the session’s cost goes | ||
| Cache writes | $0.80 | $1.20 |
| Cache reads | $0.50 | $0.60 |
| Uncached input | $0.20 | $0.30 |
| Output | $0.30 | $0.75 |
- caching saves on the session with Qwen3.8-Max (66%)
- $3.50
- caching saves on the session with Kimi K3 (65%)
- $5.40
Open weights: Kimi K3 has them, Qwen3.8-Max does not
These two flagships take opposite approaches to weights. Moonshot publishes open weights for Kimi K3 under its own Kimi K3 license and describes it as a 2.8 trillion parameter model. Qwen3.8-Max is a closed API model, and Qwen publishes open weights for its base variant, Qwen3.8-2.4T-A95B, rather than for the Max model itself.
In practice, Kimi K3 can be self-hosted or served by several OpenRouter providers, whose prices can differ from Moonshot's. Qwen3.8-Max is reached through Qwen Cloud, Alibaba Cloud Model Studio, OpenRouter, or OpenCode. This page does not estimate what self-hosting Kimi K3 would cost, and the figures below use each maker's own API price.
Qwen3.8-Max is cheaper on every line
Qwen3.8-Max lists $2 per million input and $6 per million output, where Kimi K3 lists $3 and $15, so input is 1.5x dearer on Kimi K3 and output 2.5x. The large one-off review costs $0.36 against $0.60, and the output-heavy generation $0.54 against $1.29.
On the agentic session, Qwen3.8-Max comes to $1.80 and Kimi K3 to $2.85, a $1.05 gap. Output is the largest piece at $0.45, then cache writes at $0.40, since both makers bill a default write at the input rate. Reads and fresh input add $0.10 each. At 110 sessions a month that is $198.00 against $313.50.
Cache hits are the closest rate: $0.25 on Qwen's implicit cache and $0.30 on Kimi K3, 17% apart. Caching saves almost the same share on both, 66% and 65% of the uncached session. The homepage cache example follows the session step by step.
How Qwen and Kimi prompt caching compare
Both makers cache automatically. Qwen's implicit cache cannot be turned off and bills written tokens as input. Kimi's lifetimes are 5 minutes by default and 1 hour on request, at $3 and $6 per million written, and a hit restarts the clock without charge.
Qwen adds an explicit mode, marked with cache_control, that lasts 5 minutes and costs $2.50 per million to create and $0.17 per million to read. Against Kimi's 5-minute cache that is cheaper on both sides: $2.50 against $3 to write and $0.17 against $0.30 to read. This comparison prices Qwen's implicit mode.
High hit rates favor both models. Moonshot cites a rate above 90% for Kimi K3 coding workloads on its own API, and every hit is billed at the lowest rate either maker charges.
Where they run, and how Qwen and Moonshot pitch them
Moonshot's model reaches more tools: Cursor, OpenCode, OpenRouter, and GitHub Copilot all offer Kimi K3, and Moonshot's own API opens after a minimum $1 top-up. Qwen3.8-Max is offered through OpenRouter and in OpenCode. Qwen says it trained its model with reinforcement learning across harnesses including Claude Code and Codex, two agents built around their own makers' models.
Both makers point at long, unattended coding. Qwen calls Qwen3.8-Max "the most capable model in the Qwen family to date" and says it built a self-evolving harness during an autonomous coding run of more than 10 days. Moonshot calls Kimi K3 its most capable flagship to date and pitches it for long engineering tasks with minimal supervision, large codebases, and terminal tools.
Output limits set them apart. Kimi K3's default of 131,072 output tokens per request is about Qwen3.8-Max's 131K cap, but Kimi can raise it to the full 1.05M window, about 8x as much. Context is close, at 1.05M and 1M. EveryToken prices both from OpenRouter's catalog when you use them through OpenRouter, and does not price direct calls to Qwen Cloud or Moonshot's API.
Prompt caching
How each maker bills cached tokens
Alibaba Qwen
Implicit caching is automatic and can't be turned off. A hit on Qwen3.8-Max costs $0.25 per million tokens, and written tokens cost ordinary input.
Explicit caching, marked with cache_control, lasts 5 minutes. On Qwen3.8-Max it costs $2.50 per million tokens to create and $0.17 per million to read. This blog prices Qwen with implicit caching.
Moonshot AI
The Kimi API caches prompt prefixes automatically, for 5 minutes by default or 1 hour on request. A cache hit on Kimi K3 costs one-tenth of the input price, and each hit restarts the cache lifetime at no charge.
Kimi bills cache writes separately: $3 per million tokens for the 5-minute cache, the same as ordinary input, and $6 for the 1-hour cache. This blog prices Kimi's writes at the 5-minute default.
Source: Kimi API: Chat pricing
Your own numbers
See what Qwen3.8-Max and Kimi K3 really cost you.
everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is Qwen3.8-Max cheaper than Kimi K3?
Yes, on every rate. The example agentic session costs $1.80 against $2.85, the uncached review $0.36 against $0.60, and output-heavy generation $0.54 against $1.29.
Which one has open weights?
Kimi K3, under Moonshot's own Kimi K3 license. Qwen3.8-Max is a closed API model, and Qwen publishes open weights for its base variant, Qwen3.8-2.4T-A95B, instead.
Can I use Qwen3.8-Max in Cursor or GitHub Copilot?
No. Both offer Kimi K3, and neither offers Qwen3.8-Max. Qwen3.8-Max is available through OpenRouter, in OpenCode, and on Qwen Cloud.
Which model charges less for cache hits?
Qwen3.8-Max: $0.25 per million with its automatic cache and $0.17 with explicit caching, against $0.30 on Kimi K3.
Sources
- Qwen Cloud: Qwen3.8-Max
- Qwen: Qwen3.8
- Alibaba Cloud Model Studio: Model pricing
- OpenRouter: Qwen3.8-Max
- OpenCode docs: Zen
- Kimi API: Chat pricing
- Kimi API: Kimi K3 quickstart
- Kimi: Kimi K3
- OpenRouter: Kimi K3
- Cursor docs: Models
- GitHub Docs: Supported AI models in Copilot
- Alibaba Cloud Model Studio: Context cache