Skip to content

Model comparison

GLM-5.3 vs GLM-5.3-Flash: is Z.ai's flagship worth 9x?

GLM-5.3-Flash costs about a ninth of GLM-5.3 at Z.ai's rates: $0.16 against $1.44 for an agentic coding session. Same limits, same caching, different jobs.

· Prices as of September 28, 2026

  • GLM-5.3

    Z.ai · Released August 14, 2026

    Z.ai's flagship for complex software engineering and long-horizon agent work, with open weights.

    GLM-5.3 facts and comparisons
  • GLM-5.3-Flash

    Z.ai · Released August 26, 2026

    Z.ai's low-cost, natively multimodal model, which Z.ai says outperforms GLM-5.2 at a tenth of the price.

    GLM-5.3-Flash facts and comparisons

The short answer

GLM-5.3-Flash costs $0.16 on the example agentic coding session against $1.44 on GLM-5.3, a 9x gap, with the same 1M context, 128K output limit, and caching rules. Choose GLM-5.3 for the work Z.ai aims its flagship at, complex software engineering and long-horizon agents; choose GLM-5.3-Flash for high-volume or visual coding work, where Z.ai pitches its multimodal Flash model.

Choose GLM-5.3 if

  • You want Z.ai's flagship, which it positions for complex software engineering and long-horizon agent work.
  • You would rather pay a flat rate: Z.ai sells GLM-5.3 with its GLM Coding Plan subscription.
  • The per-session difference, $1.28 on the example session, is small next to the value of the task, so price does not decide it.

Choose GLM-5.3-Flash if

  • You send many requests, where the monthly gap adds up: 110 sessions cost $17.60 on GLM-5.3-Flash against $158.40.
  • Your coding work is visual: Z.ai says GLM-5.3-Flash looks at interfaces and rendered results to test and improve its work.
  • You want open weights under the MIT license rather than Z.ai's own GLM-5.3 license.
  • You want a model many developers already use: GLM-5.3-Flash was the most-used model for programming on OpenRouter over the week before September 28, 2026, summed across nine languages.

Side by side

Specs and prices

FactGLM-5.3GLM-5.3-Flash
MakerZ.aiZ.ai
API model idglm-5.3glm-5.3-flash
ReleasedAugust 14, 2026August 26, 2026
StatusCurrentCurrent
Context window1M tokens1M tokens
Max output128K tokens128K tokens
Open weightsYesYes
Input, per 1M tokens$1.40$0.15
Cache hit, per 1M$0.26$0.03
Cache write, per 1M$1.40 (same as input)$0.15 (same as input)
Output, per 1M tokens$4.40$0.50
Runs inOpenCode and OpenRouterOpenCode and OpenRouter

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadGLM-5.3GLM-5.3-Flash
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$1.44$0.16
Large one-off review, 150K input with no cache hits, 10K output$0.25$0.03
Output-heavy generation, 30K input, 80K output$0.39$0.04
A month of sessions, 110 sessions: 5 a day, 22 working days$158.40$17.60
Where the session’s cost goes
Cache writes$0.56$0.06
Cache reads$0.52$0.06
Uncached input$0.14$0.02
Output$0.22$0.03
caching saves on the session with GLM-5.3 (61%)
$2.28
caching saves on the session with GLM-5.3-Flash (60%)
$0.24

How much cheaper is GLM-5.3-Flash than GLM-5.3?

At Z.ai's own rates, GLM-5.3-Flash charges $0.15 per million input tokens, $0.03 per million cache hits, and $0.50 per million output tokens. GLM-5.3 charges $1.40, $0.26, and $4.40. Every rate is between 8.7x and 9.3x higher on the flagship.

The example workloads follow suit. The agentic coding session costs $1.44 on GLM-5.3 and $0.16 on GLM-5.3-Flash, 9x apart. The large one-off review costs $0.25 against $0.03, and the output-heavy generation $0.39 against $0.04. At 110 sessions a month, the session totals $158.40 against $17.60, a difference of $140.80.

Flash's figures are small enough that rounding to the cent shows. Its session lines, $0.06 of cache writes, $0.06 of cache reads, $0.02 of fresh input, and $0.03 of output, are each rounded, so ratios on a single Flash workload are approximate. The monthly totals are the steadier comparison.

Same limits and caching rules at two price points

The limits match. Both models take 1M tokens of context and write up to 128K tokens of output, both are offered through OpenRouter and in OpenCode, and both use Z.ai's automatic caching, which needs no configuration and carries no fee for writing the cache.

The cache discount is nearly the same share of input: a hit is 18.6% of input on GLM-5.3 and 20% on GLM-5.3-Flash. Caching therefore saves a similar share on each, 61% and 60% of the uncached session, which in dollars is $2.28 on GLM-5.3 and $0.24 on Flash. Z.ai says cached input storage is free for a limited time.

Because the rules match, moving work between the two changes what a session costs, not how it should be structured. The cache example on our homepage shows the session both sets of figures come from.

What Z.ai says each GLM model is for

Z.ai calls GLM-5.3 its "latest flagship model, delivering comprehensive advancements in complex software engineering and agent capabilities," and positions it for long-horizon agent work. Reasoning is always on, at low, high, or max, and Z.ai serves it on OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages endpoints.

GLM-5.3-Flash is Z.ai's low-cost, natively multimodal model, released on August 26, 2026, twelve days after GLM-5.3. Z.ai says it outperforms GLM-5.2 at a tenth of the price, and describes native multimodal visual coding, in which the model looks at interfaces and rendered results to test and improve its work. Z.ai also credits a hybrid of sparse and linear attention with cutting attention compute and KV cache size.

Flash already has wide use. By OpenRouter's count, summed across nine languages, Flash was the most-used model for programming over the week before September 28, 2026. That is a usage figure rather than a quality claim, but it shows how many developers already point code at it.

Licenses, self-hosting, and tracking both

Both models ship open weights, under different terms. GLM-5.3-Flash uses the MIT license. GLM-5.3 uses Z.ai's own GLM-5.3 license, so read its terms before building on the weights. This page does not estimate hosting costs for either.

The prices here are Z.ai's own API rates. OpenRouter sends GLM requests to one of several providers, whose prices can differ. EveryToken prices both GLM models from OpenRouter's catalog when you use them through OpenRouter, so a mix of flagship and Flash requests shows up side by side; it does not price calls made directly to Z.ai's API.

Prompt caching

How Z.ai bills cached tokens

Z.ai

Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.

A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.

Source: Z.ai docs: Context caching

Your own numbers

See what GLM-5.3 and GLM-5.3-Flash really cost you.

everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Does GLM-5.3 really cost nine times as much as GLM-5.3-Flash?

Close to it. Every list rate is 8.7x to 9.3x higher on GLM-5.3, and the example agentic session costs $1.44 against $0.16, a 9x gap. At 110 sessions a month that is $158.40 against $17.60.

Do GLM-5.3 and GLM-5.3-Flash have the same limits?

Yes. GLM-5.3 and GLM-5.3-Flash both take 1M tokens of context and write up to 128K tokens of output. Price, license, and the work Z.ai aims each at are what separate them.

Which license does each GLM model use?

GLM-5.3-Flash is released under the MIT license. GLM-5.3 uses Z.ai's own GLM-5.3 license. Both GLM models have open weights, so either can be self-hosted within its license terms.

Do both models cache the same way?

Yes. Z.ai caches repeated context automatically for both and lists no fee for writing the cache. A hit costs $0.26 per million on GLM-5.3 and $0.03 on GLM-5.3-Flash.

  • GLM-5.3-Flash vs Claude Sonnet 5

    GLM-5.3-Flash costs $0.16 for a cached coding session that costs $2.40 on Claude Sonnet 5. What drives a 15x gap, and what price alone can't tell you.

  • DeepSeek-V4-Pro vs GLM-5.3

    DeepSeek-V4-Pro and GLM-5.3 list input and output within 11% of each other, but a DeepSeek cache hit costs $0.044 against $0.26. What that does to a session.

  • DeepSeek-V4.1-Flash vs GLM-5.3-Flash

    GLM-5.3-Flash lists half DeepSeek-V4.1-Flash's input price, but GLM's cache hits cost 5x as much. A cached coding session: $0.16 against $0.22 at peak.

  • GLM-5.3 vs Claude Sonnet 5

    GLM-5.3 undercuts Claude Sonnet 5 by 40% on an agentic coding session, $1.44 against $2.40. Most of it comes from Anthropic's cache-write premium.

  • GLM-5.3 vs Claude Opus 5.5

    GLM-5.3 costs about a third of Claude Opus 5.5 on an agentic coding session, $1.44 against $4.40, yet its cache hits cost more. Where the gap comes from.

  • GLM-5.3 vs Gemini 3.1 Pro Preview

    Neither Z.ai nor Google charges extra to write the cache, so GLM-5.3 and Gemini 3.1 Pro Preview split on output: $4.40 against $12 per million tokens.