Skip to content

Model comparison

Kimi K3 vs GLM-5.3: open-weight flagships compared on cost

Kimi K3 and GLM-5.3 are both open-weight flagships. A coding session costs $2.85 on Kimi K3 and $1.44 on GLM-5.3, yet their cache hits nearly match.

· Prices as of September 28, 2026

  • Kimi K3

    Moonshot AI · Released July 16, 2026

    Moonshot's flagship and the largest open model it has released, aimed at long-horizon coding and knowledge work.

    Kimi K3 facts and comparisons
  • GLM-5.3

    Z.ai · Released August 14, 2026

    Z.ai's flagship for complex software engineering and long-horizon agent work, with open weights.

    GLM-5.3 facts and comparisons

The short answer

GLM-5.3 costs about half as much as Kimi K3 on the example agentic coding session, $1.44 against $2.85, while output-heavy work costs 3.3x as much on Kimi K3, although the two charge almost the same for a cache hit. Pick Kimi K3 if you want it in Cursor or GitHub Copilot or need responses beyond 128K tokens; pick GLM-5.3 if price leads and OpenRouter or OpenCode fit your workflow.

Choose Kimi K3 if

  • You pick models in Cursor or GitHub Copilot, which offer Kimi K3 and do not offer GLM-5.3.
  • You need responses longer than GLM-5.3's 128K limit, since Kimi K3 can raise its output to its full 1.05M window.
  • You want Moonshot's largest open model, which it calls its most capable flagship to date, with 2.8 trillion parameters.

Choose GLM-5.3 if

  • Price leads: GLM-5.3 charges $1.40 input and $4.40 output per million tokens, against $3 and $15 on Kimi K3.
  • Your work is output-heavy, where the gap is widest: the example generation costs $0.39 on GLM-5.3 against $1.29.
  • You want a subscription option, since Z.ai sells GLM-5.3 with its GLM Coding Plan.

Side by side

Specs and prices

FactKimi K3GLM-5.3
MakerMoonshot AIZ.ai
API model idkimi-k3glm-5.3
ReleasedJuly 16, 2026August 14, 2026
StatusCurrentCurrent
Context window1.05M tokens1M tokens
Max output1.05M tokens128K tokens
Open weightsYesYes
Input, per 1M tokens$3$1.40
Cache hit, per 1M$0.30$0.26
Cache write, per 1M$3 (same as input)$1.40 (same as input)
Output, per 1M tokens$15$4.40
Runs inCursor, OpenCode, OpenRouter, and GitHub CopilotOpenCode and OpenRouter

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Kimi K3: A 1-hour cache write costs $6 per million tokens; this table uses the 5-minute default.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadKimi K3GLM-5.3
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$2.85$1.44
Large one-off review, 150K input with no cache hits, 10K output$0.60$0.25
Output-heavy generation, 30K input, 80K output$1.29$0.39
A month of sessions, 110 sessions: 5 a day, 22 working days$313.50$158.40
Where the session’s cost goes
Cache writes$1.20$0.56
Cache reads$0.60$0.52
Uncached input$0.30$0.14
Output$0.75$0.22
caching saves on the session with Kimi K3 (65%)
$5.40
caching saves on the session with GLM-5.3 (61%)
$2.28

Why Kimi K3 costs about twice as much as GLM-5.3

Kimi K3 lists $3 per million input tokens and $15 per million output tokens. GLM-5.3 lists $1.40 and $4.40, so input is 2.1x dearer on Kimi K3 and output 3.4x. Uncached work shows both: the large one-off review costs $0.60 against $0.25, and the output-heavy generation $1.29 against $0.39.

The agentic session lands at 2x, $2.85 on Kimi K3 against $1.44 on GLM-5.3, a $1.41 difference. Cache writes account for $0.64 of it, since both makers bill a default write at the input rate and Kimi's input is dearer. Output adds $0.53 and fresh input $0.16.

Over 110 sessions a month, that comes to $313.50 against $158.40, or $155.10 apart. These are API-equivalent estimates at Moonshot's and Z.ai's own API rates.

Their cache hits cost almost the same

Reads are the exception. A Kimi K3 cache hit costs $0.30 per million, one-tenth of its input price. A GLM-5.3 hit costs $0.26, or 18.6% of input. Kimi's deeper discount nearly cancels its higher input price and leaves the two hits 13% apart.

In the session, 2M cached tokens cost $0.60 on Kimi K3 and $0.52 on GLM-5.3, only $0.08 apart. The larger the share of a session that comes from the cache, the smaller the relative gap, and the more fresh context and output it produces, the larger. Caching saves 65% on Kimi K3 and 61% on GLM-5.3 against the uncached cost, and the homepage cache example shows the session in detail.

The two caches behave differently. Kimi keeps a prefix for 5 minutes by default or 1 hour on request, bills the 1-hour write at $6 per million, and restarts the lifetime free on every hit. For GLM-5.3, Z.ai caches repeated context with no write fee and says cached input storage is free for a limited time. Moonshot's own figure for coding on its API is a hit rate above 90%.

Open weights, licenses, and where to run them

Both models ship open weights, each under its maker's own license: Moonshot's Kimi K3 license and Z.ai's GLM-5.3 license. Kimi K3 and GLM-5.3 can each be self-hosted, and this page does not estimate what that would cost. The Kimi K3 and GLM-5.3 prices here are the makers' own API rates, and OpenRouter providers serving the weights can charge differently.

Kimi K3 is offered in Cursor, OpenCode, OpenRouter, and GitHub Copilot, and on Moonshot's API after a minimum $1 top-up. On the GLM side, OpenRouter and OpenCode carry the model, and Z.ai's own API takes requests in OpenAI Chat Completions, OpenAI Responses, or Anthropic Messages form. Z.ai also sells GLM-5.3 with its GLM Coding Plan.

EveryToken prices Kimi K3 and GLM-5.3 from OpenRouter's catalog when you use them through OpenRouter. It does not price either one called directly on Moonshot's or Z.ai's own API.

Output limits and how each maker pitches its flagship

Kimi K3's output can reach its full 1.05M context window, with a default of 131,072 tokens per request. GLM-5.3 writes up to 128K. Context is close, 1.05M against 1M.

Moonshot calls Kimi K3 "Kimi's most capable flagship model to date, with 2.8 trillion parameters," and pitches it at long engineering tasks with minimal supervision, large codebases, and terminal tools, plus frontend and game work that combines code with visual reasoning from screenshots. Z.ai calls GLM-5.3 its latest flagship, "delivering comprehensive advancements in complex software engineering and agent capabilities," with reasoning always on at low, high, or max.

Both are pitched at long agentic coding, so the makers' words will not separate them. Tokenizers differ, reasoning settings change output length, and output is where these two differ most in price, so run a representative task on each before choosing.

Prompt caching

How each maker bills cached tokens

Moonshot AI

The Kimi API caches prompt prefixes automatically, for 5 minutes by default or 1 hour on request. A cache hit on Kimi K3 costs one-tenth of the input price, and each hit restarts the cache lifetime at no charge.

Kimi bills cache writes separately: $3 per million tokens for the 5-minute cache, the same as ordinary input, and $6 for the 1-hour cache. This blog prices Kimi's writes at the 5-minute default.

Source: Kimi API: Chat pricing

Z.ai

Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.

A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.

Source: Z.ai docs: Context caching

Your own numbers

See what Kimi K3 and GLM-5.3 really cost you.

everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is GLM-5.3 cheaper than Kimi K3?

Yes, on every rate. The example agentic session costs $1.44 on GLM-5.3 against $2.85 on Kimi K3, and output-heavy generation $0.39 against $1.29. Cache hits are closest, at $0.26 against $0.30 per million.

Do Kimi K3 and GLM-5.3 have open weights?

Yes. Moonshot publishes Kimi K3 under its own Kimi K3 license, and Z.ai publishes GLM-5.3 under its own GLM-5.3 license. Read each license before self-hosting or building on the weights.

Which one can I use in GitHub Copilot?

Kimi K3. GitHub Copilot and Cursor both offer it, and neither offers GLM-5.3. Both models are available through OpenRouter and in OpenCode.

Does Kimi charge for cache writes?

Yes, as a separate line: Kimi K3 writes cost $3 per million for the 5-minute cache, equal to input, and $6 for the 1-hour cache. Z.ai lists no fee for writing the cache, so GLM-5.3 writes cost its $1.40 input rate.

  • GLM-5.3 vs GLM-5.3-Flash

    GLM-5.3-Flash costs about a ninth of GLM-5.3 at Z.ai's rates: $0.16 against $1.44 for an agentic coding session. Same limits, same caching, different jobs.

  • DeepSeek-V4-Pro vs GLM-5.3

    DeepSeek-V4-Pro and GLM-5.3 list input and output within 11% of each other, but a DeepSeek cache hit costs $0.044 against $0.26. What that does to a session.

  • GLM-5.3 vs Claude Sonnet 5

    GLM-5.3 undercuts Claude Sonnet 5 by 40% on an agentic coding session, $1.44 against $2.40. Most of it comes from Anthropic's cache-write premium.

  • GLM-5.3 vs Claude Opus 5.5

    GLM-5.3 costs about a third of Claude Opus 5.5 on an agentic coding session, $1.44 against $4.40, yet its cache hits cost more. Where the gap comes from.

  • GLM-5.3 vs Gemini 3.1 Pro Preview

    Neither Z.ai nor Google charges extra to write the cache, so GLM-5.3 and Gemini 3.1 Pro Preview split on output: $4.40 against $12 per million tokens.

  • GLM-5.3 vs GPT-6 Sol

    GLM-5.3 costs 31% less than GPT-6 Sol on an agentic coding session, $1.44 against $2.10. How OpenAI's write fee and 272K rule compare with Z.ai's rates.