Model comparison
Kimi K3 vs GLM-5.3: open-weight flagships compared on cost
Kimi K3 and GLM-5.3 are both open-weight flagships. A coding session costs $2.85 on Kimi K3 and $1.44 on GLM-5.3, yet their cache hits nearly match.
· Prices as of September 28, 2026
Kimi K3
Moonshot AI · Released July 16, 2026
Moonshot's flagship and the largest open model it has released, aimed at long-horizon coding and knowledge work.
Kimi K3 facts and comparisonsGLM-5.3
Z.ai · Released August 14, 2026
Z.ai's flagship for complex software engineering and long-horizon agent work, with open weights.
GLM-5.3 facts and comparisons
The short answer
GLM-5.3 costs about half as much as Kimi K3 on the example agentic coding session, $1.44 against $2.85, while output-heavy work costs 3.3x as much on Kimi K3, although the two charge almost the same for a cache hit. Pick Kimi K3 if you want it in Cursor or GitHub Copilot or need responses beyond 128K tokens; pick GLM-5.3 if price leads and OpenRouter or OpenCode fit your workflow.
Choose Kimi K3 if
- You pick models in Cursor or GitHub Copilot, which offer Kimi K3 and do not offer GLM-5.3.
- You need responses longer than GLM-5.3's 128K limit, since Kimi K3 can raise its output to its full 1.05M window.
- You want Moonshot's largest open model, which it calls its most capable flagship to date, with 2.8 trillion parameters.
Choose GLM-5.3 if
- Price leads: GLM-5.3 charges $1.40 input and $4.40 output per million tokens, against $3 and $15 on Kimi K3.
- Your work is output-heavy, where the gap is widest: the example generation costs $0.39 on GLM-5.3 against $1.29.
- You want a subscription option, since Z.ai sells GLM-5.3 with its GLM Coding Plan.
Side by side
Specs and prices
| Fact | Kimi K3 | GLM-5.3 |
|---|---|---|
| Maker | Moonshot AI | Z.ai |
| API model id | kimi-k3 | glm-5.3 |
| Released | July 16, 2026 | August 14, 2026 |
| Status | Current | Current |
| Context window | 1.05M tokens | 1M tokens |
| Max output | 1.05M tokens | 128K tokens |
| Open weights | Yes | Yes |
| Input, per 1M tokens | $3 | $1.40 |
| Cache hit, per 1M | $0.30 | $0.26 |
| Cache write, per 1M | $3 (same as input) | $1.40 (same as input) |
| Output, per 1M tokens | $15 | $4.40 |
| Runs in | Cursor, OpenCode, OpenRouter, and GitHub Copilot | OpenCode and OpenRouter |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Kimi K3: A 1-hour cache write costs $6 per million tokens; this table uses the 5-minute default.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Kimi K3 | GLM-5.3 |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $2.85 | $1.44 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.60 | $0.25 |
| Output-heavy generation, 30K input, 80K output | $1.29 | $0.39 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $313.50 | $158.40 |
| Where the session’s cost goes | ||
| Cache writes | $1.20 | $0.56 |
| Cache reads | $0.60 | $0.52 |
| Uncached input | $0.30 | $0.14 |
| Output | $0.75 | $0.22 |
- caching saves on the session with Kimi K3 (65%)
- $5.40
- caching saves on the session with GLM-5.3 (61%)
- $2.28
Why Kimi K3 costs about twice as much as GLM-5.3
Kimi K3 lists $3 per million input tokens and $15 per million output tokens. GLM-5.3 lists $1.40 and $4.40, so input is 2.1x dearer on Kimi K3 and output 3.4x. Uncached work shows both: the large one-off review costs $0.60 against $0.25, and the output-heavy generation $1.29 against $0.39.
The agentic session lands at 2x, $2.85 on Kimi K3 against $1.44 on GLM-5.3, a $1.41 difference. Cache writes account for $0.64 of it, since both makers bill a default write at the input rate and Kimi's input is dearer. Output adds $0.53 and fresh input $0.16.
Over 110 sessions a month, that comes to $313.50 against $158.40, or $155.10 apart. These are API-equivalent estimates at Moonshot's and Z.ai's own API rates.
Their cache hits cost almost the same
Reads are the exception. A Kimi K3 cache hit costs $0.30 per million, one-tenth of its input price. A GLM-5.3 hit costs $0.26, or 18.6% of input. Kimi's deeper discount nearly cancels its higher input price and leaves the two hits 13% apart.
In the session, 2M cached tokens cost $0.60 on Kimi K3 and $0.52 on GLM-5.3, only $0.08 apart. The larger the share of a session that comes from the cache, the smaller the relative gap, and the more fresh context and output it produces, the larger. Caching saves 65% on Kimi K3 and 61% on GLM-5.3 against the uncached cost, and the homepage cache example shows the session in detail.
The two caches behave differently. Kimi keeps a prefix for 5 minutes by default or 1 hour on request, bills the 1-hour write at $6 per million, and restarts the lifetime free on every hit. For GLM-5.3, Z.ai caches repeated context with no write fee and says cached input storage is free for a limited time. Moonshot's own figure for coding on its API is a hit rate above 90%.
Open weights, licenses, and where to run them
Both models ship open weights, each under its maker's own license: Moonshot's Kimi K3 license and Z.ai's GLM-5.3 license. Kimi K3 and GLM-5.3 can each be self-hosted, and this page does not estimate what that would cost. The Kimi K3 and GLM-5.3 prices here are the makers' own API rates, and OpenRouter providers serving the weights can charge differently.
Kimi K3 is offered in Cursor, OpenCode, OpenRouter, and GitHub Copilot, and on Moonshot's API after a minimum $1 top-up. On the GLM side, OpenRouter and OpenCode carry the model, and Z.ai's own API takes requests in OpenAI Chat Completions, OpenAI Responses, or Anthropic Messages form. Z.ai also sells GLM-5.3 with its GLM Coding Plan.
EveryToken prices Kimi K3 and GLM-5.3 from OpenRouter's catalog when you use them through OpenRouter. It does not price either one called directly on Moonshot's or Z.ai's own API.
Output limits and how each maker pitches its flagship
Kimi K3's output can reach its full 1.05M context window, with a default of 131,072 tokens per request. GLM-5.3 writes up to 128K. Context is close, 1.05M against 1M.
Moonshot calls Kimi K3 "Kimi's most capable flagship model to date, with 2.8 trillion parameters," and pitches it at long engineering tasks with minimal supervision, large codebases, and terminal tools, plus frontend and game work that combines code with visual reasoning from screenshots. Z.ai calls GLM-5.3 its latest flagship, "delivering comprehensive advancements in complex software engineering and agent capabilities," with reasoning always on at low, high, or max.
Both are pitched at long agentic coding, so the makers' words will not separate them. Tokenizers differ, reasoning settings change output length, and output is where these two differ most in price, so run a representative task on each before choosing.
Prompt caching
How each maker bills cached tokens
Moonshot AI
The Kimi API caches prompt prefixes automatically, for 5 minutes by default or 1 hour on request. A cache hit on Kimi K3 costs one-tenth of the input price, and each hit restarts the cache lifetime at no charge.
Kimi bills cache writes separately: $3 per million tokens for the 5-minute cache, the same as ordinary input, and $6 for the 1-hour cache. This blog prices Kimi's writes at the 5-minute default.
Source: Kimi API: Chat pricing
Z.ai
Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.
A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.
Source: Z.ai docs: Context caching
Your own numbers
See what Kimi K3 and GLM-5.3 really cost you.
everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is GLM-5.3 cheaper than Kimi K3?
Yes, on every rate. The example agentic session costs $1.44 on GLM-5.3 against $2.85 on Kimi K3, and output-heavy generation $0.39 against $1.29. Cache hits are closest, at $0.26 against $0.30 per million.
Do Kimi K3 and GLM-5.3 have open weights?
Yes. Moonshot publishes Kimi K3 under its own Kimi K3 license, and Z.ai publishes GLM-5.3 under its own GLM-5.3 license. Read each license before self-hosting or building on the weights.
Which one can I use in GitHub Copilot?
Kimi K3. GitHub Copilot and Cursor both offer it, and neither offers GLM-5.3. Both models are available through OpenRouter and in OpenCode.
Does Kimi charge for cache writes?
Yes, as a separate line: Kimi K3 writes cost $3 per million for the 5-minute cache, equal to input, and $6 for the 1-hour cache. Z.ai lists no fee for writing the cache, so GLM-5.3 writes cost its $1.40 input rate.