Skip to content

Model comparison

Qwen3.8-Max vs Kimi K3: a closed API against open weights

Qwen3.8-Max undercuts Kimi K3 on every rate, $1.80 against $2.85 for an agentic coding session. Kimi K3 has open weights; the Qwen3.8-Max API model does not.

· Prices as of September 28, 2026

  • Qwen3.8-Max

    Alibaba Qwen · Released August 2, 2026

    Qwen's most capable model, built for long autonomous coding and professional work.

    Qwen3.8-Max facts and comparisons
  • Kimi K3

    Moonshot AI · Released July 16, 2026

    Moonshot's flagship and the largest open model it has released, aimed at long-horizon coding and knowledge work.

    Kimi K3 facts and comparisons

The short answer

Qwen3.8-Max is cheaper on every rate, and the example agentic coding session costs $1.80 on it against $2.85 on Kimi K3, 37% less. Pick Kimi K3 if you want open weights, access in Cursor or GitHub Copilot, or responses longer than 131K tokens; pick Qwen3.8-Max if a hosted API is all you need and you want the lower price through OpenRouter, OpenCode, or Qwen Cloud.

Choose Qwen3.8-Max if

  • You want the lower price everywhere: $2 input, $0.25 cache hits, and $6 output per million, against $3, $0.30, and $15.
  • Your work is output-heavy, where the gap is widest: the example generation costs $0.54 on Qwen3.8-Max against $1.29.
  • Your tool can mark cache breakpoints, where Qwen's explicit cache reads at $0.17 per million, well under Kimi K3's $0.30.

Choose Kimi K3 if

  • You need open weights for the exact model you run: Moonshot publishes Kimi K3's, while Qwen keeps the Max API model closed.
  • You choose models in Cursor or GitHub Copilot, which offer Kimi K3 and do not offer Qwen3.8-Max.
  • You need responses longer than 131K tokens, since Kimi K3 can raise output to its full 1.05M window.

Side by side

Specs and prices

FactQwen3.8-MaxKimi K3
MakerAlibaba QwenMoonshot AI
API model idqwen3.8-maxkimi-k3
ReleasedAugust 2, 2026July 16, 2026
StatusCurrentCurrent
Context window1M tokens1.05M tokens
Max output131K tokens1.05M tokens
Open weightsNoYes
Input, per 1M tokens$2$3
Cache hit, per 1M$0.25$0.30
Cache write, per 1M$2 (same as input)$3 (same as input)
Output, per 1M tokens$6$15
Runs inOpenCode and OpenRouterCursor, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Qwen3.8-Max: Qwen Cloud prices. Alibaba Cloud Model Studio's international scope lists the same input and output rates. Kimi K3: A 1-hour cache write costs $6 per million tokens; this table uses the 5-minute default.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadQwen3.8-MaxKimi K3
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$1.80$2.85
Large one-off review, 150K input with no cache hits, 10K output$0.36$0.60
Output-heavy generation, 30K input, 80K output$0.54$1.29
A month of sessions, 110 sessions: 5 a day, 22 working days$198.00$313.50
Where the session’s cost goes
Cache writes$0.80$1.20
Cache reads$0.50$0.60
Uncached input$0.20$0.30
Output$0.30$0.75
caching saves on the session with Qwen3.8-Max (66%)
$3.50
caching saves on the session with Kimi K3 (65%)
$5.40

Open weights: Kimi K3 has them, Qwen3.8-Max does not

These two flagships take opposite approaches to weights. Moonshot publishes open weights for Kimi K3 under its own Kimi K3 license and describes it as a 2.8 trillion parameter model. Qwen3.8-Max is a closed API model, and Qwen publishes open weights for its base variant, Qwen3.8-2.4T-A95B, rather than for the Max model itself.

In practice, Kimi K3 can be self-hosted or served by several OpenRouter providers, whose prices can differ from Moonshot's. Qwen3.8-Max is reached through Qwen Cloud, Alibaba Cloud Model Studio, OpenRouter, or OpenCode. This page does not estimate what self-hosting Kimi K3 would cost, and the figures below use each maker's own API price.

Qwen3.8-Max is cheaper on every line

Qwen3.8-Max lists $2 per million input and $6 per million output, where Kimi K3 lists $3 and $15, so input is 1.5x dearer on Kimi K3 and output 2.5x. The large one-off review costs $0.36 against $0.60, and the output-heavy generation $0.54 against $1.29.

On the agentic session, Qwen3.8-Max comes to $1.80 and Kimi K3 to $2.85, a $1.05 gap. Output is the largest piece at $0.45, then cache writes at $0.40, since both makers bill a default write at the input rate. Reads and fresh input add $0.10 each. At 110 sessions a month that is $198.00 against $313.50.

Cache hits are the closest rate: $0.25 on Qwen's implicit cache and $0.30 on Kimi K3, 17% apart. Caching saves almost the same share on both, 66% and 65% of the uncached session. The homepage cache example follows the session step by step.

How Qwen and Kimi prompt caching compare

Both makers cache automatically. Qwen's implicit cache cannot be turned off and bills written tokens as input. Kimi's lifetimes are 5 minutes by default and 1 hour on request, at $3 and $6 per million written, and a hit restarts the clock without charge.

Qwen adds an explicit mode, marked with cache_control, that lasts 5 minutes and costs $2.50 per million to create and $0.17 per million to read. Against Kimi's 5-minute cache that is cheaper on both sides: $2.50 against $3 to write and $0.17 against $0.30 to read. This comparison prices Qwen's implicit mode.

High hit rates favor both models. Moonshot cites a rate above 90% for Kimi K3 coding workloads on its own API, and every hit is billed at the lowest rate either maker charges.

Where they run, and how Qwen and Moonshot pitch them

Moonshot's model reaches more tools: Cursor, OpenCode, OpenRouter, and GitHub Copilot all offer Kimi K3, and Moonshot's own API opens after a minimum $1 top-up. Qwen3.8-Max is offered through OpenRouter and in OpenCode. Qwen says it trained its model with reinforcement learning across harnesses including Claude Code and Codex, two agents built around their own makers' models.

Both makers point at long, unattended coding. Qwen calls Qwen3.8-Max "the most capable model in the Qwen family to date" and says it built a self-evolving harness during an autonomous coding run of more than 10 days. Moonshot calls Kimi K3 its most capable flagship to date and pitches it for long engineering tasks with minimal supervision, large codebases, and terminal tools.

Output limits set them apart. Kimi K3's default of 131,072 output tokens per request is about Qwen3.8-Max's 131K cap, but Kimi can raise it to the full 1.05M window, about 8x as much. Context is close, at 1.05M and 1M. EveryToken prices both from OpenRouter's catalog when you use them through OpenRouter, and does not price direct calls to Qwen Cloud or Moonshot's API.

Prompt caching

How each maker bills cached tokens

Alibaba Qwen

Implicit caching is automatic and can't be turned off. A hit on Qwen3.8-Max costs $0.25 per million tokens, and written tokens cost ordinary input.

Explicit caching, marked with cache_control, lasts 5 minutes. On Qwen3.8-Max it costs $2.50 per million tokens to create and $0.17 per million to read. This blog prices Qwen with implicit caching.

Source: Alibaba Cloud Model Studio: Context cache

Moonshot AI

The Kimi API caches prompt prefixes automatically, for 5 minutes by default or 1 hour on request. A cache hit on Kimi K3 costs one-tenth of the input price, and each hit restarts the cache lifetime at no charge.

Kimi bills cache writes separately: $3 per million tokens for the 5-minute cache, the same as ordinary input, and $6 for the 1-hour cache. This blog prices Kimi's writes at the 5-minute default.

Source: Kimi API: Chat pricing

Your own numbers

See what Qwen3.8-Max and Kimi K3 really cost you.

everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is Qwen3.8-Max cheaper than Kimi K3?

Yes, on every rate. The example agentic session costs $1.80 against $2.85, the uncached review $0.36 against $0.60, and output-heavy generation $0.54 against $1.29.

Which one has open weights?

Kimi K3, under Moonshot's own Kimi K3 license. Qwen3.8-Max is a closed API model, and Qwen publishes open weights for its base variant, Qwen3.8-2.4T-A95B, instead.

Can I use Qwen3.8-Max in Cursor or GitHub Copilot?

No. Both offer Kimi K3, and neither offers Qwen3.8-Max. Qwen3.8-Max is available through OpenRouter, in OpenCode, and on Qwen Cloud.

Which model charges less for cache hits?

Qwen3.8-Max: $0.25 per million with its automatic cache and $0.17 with explicit caching, against $0.30 on Kimi K3.

  • Grok 4.7 vs Kimi K3

    Grok 4.7 and Kimi K3 both run in Cursor and GitHub Copilot. Grok 4.7 costs less on every workload, while Kimi K3 adds open weights and a 1.05M window.

  • Kimi K3 vs GPT-6 Sol

    GPT-6 Sol undercuts Kimi K3 on every rate, and an agentic coding session costs $2.10 against $2.85. Both list 1.05M context, with different limits inside.

  • Kimi K3 vs Claude Opus 5.5

    Kimi K3 lists 25% below Claude Opus 5.5 on input and output, and the gap widens to 35% on a cached coding session. Cache writes, not hits, explain it.

  • Kimi K3 vs Claude Sonnet 5

    Kimi K3 lists 1.5x the rates of Claude Sonnet 5, yet an agentic coding session costs just $0.45 more on it, $2.85 against $2.40. Cache writes explain why.

  • Kimi K3 vs DeepSeek-V4-Pro

    An agentic coding session costs $2.85 on Kimi K3 and $0.95 on DeepSeek-V4-Pro at peak rates, and DeepSeek's off-peak hours cost 50% less again.

  • Kimi K3 vs GLM-5.3

    Kimi K3 and GLM-5.3 are both open-weight flagships. A coding session costs $2.85 on Kimi K3 and $1.44 on GLM-5.3, yet their cache hits nearly match.