Model comparison
GLM-5.3 vs GLM-5.3-Flash: is Z.ai's flagship worth 9x?
GLM-5.3-Flash costs about a ninth of GLM-5.3 at Z.ai's rates: $0.16 against $1.44 for an agentic coding session. Same limits, same caching, different jobs.
· Prices as of September 28, 2026
GLM-5.3
Z.ai · Released August 14, 2026
Z.ai's flagship for complex software engineering and long-horizon agent work, with open weights.
GLM-5.3 facts and comparisonsGLM-5.3-Flash
Z.ai · Released August 26, 2026
Z.ai's low-cost, natively multimodal model, which Z.ai says outperforms GLM-5.2 at a tenth of the price.
GLM-5.3-Flash facts and comparisons
The short answer
GLM-5.3-Flash costs $0.16 on the example agentic coding session against $1.44 on GLM-5.3, a 9x gap, with the same 1M context, 128K output limit, and caching rules. Choose GLM-5.3 for the work Z.ai aims its flagship at, complex software engineering and long-horizon agents; choose GLM-5.3-Flash for high-volume or visual coding work, where Z.ai pitches its multimodal Flash model.
Choose GLM-5.3 if
- You want Z.ai's flagship, which it positions for complex software engineering and long-horizon agent work.
- You would rather pay a flat rate: Z.ai sells GLM-5.3 with its GLM Coding Plan subscription.
- The per-session difference, $1.28 on the example session, is small next to the value of the task, so price does not decide it.
Choose GLM-5.3-Flash if
- You send many requests, where the monthly gap adds up: 110 sessions cost $17.60 on GLM-5.3-Flash against $158.40.
- Your coding work is visual: Z.ai says GLM-5.3-Flash looks at interfaces and rendered results to test and improve its work.
- You want open weights under the MIT license rather than Z.ai's own GLM-5.3 license.
- You want a model many developers already use: GLM-5.3-Flash was the most-used model for programming on OpenRouter over the week before September 28, 2026, summed across nine languages.
Side by side
Specs and prices
| Fact | GLM-5.3 | GLM-5.3-Flash |
|---|---|---|
| Maker | Z.ai | Z.ai |
| API model id | glm-5.3 | glm-5.3-flash |
| Released | August 14, 2026 | August 26, 2026 |
| Status | Current | Current |
| Context window | 1M tokens | 1M tokens |
| Max output | 128K tokens | 128K tokens |
| Open weights | Yes | Yes |
| Input, per 1M tokens | $1.40 | $0.15 |
| Cache hit, per 1M | $0.26 | $0.03 |
| Cache write, per 1M | $1.40 (same as input) | $0.15 (same as input) |
| Output, per 1M tokens | $4.40 | $0.50 |
| Runs in | OpenCode and OpenRouter | OpenCode and OpenRouter |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | GLM-5.3 | GLM-5.3-Flash |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $1.44 | $0.16 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.25 | $0.03 |
| Output-heavy generation, 30K input, 80K output | $0.39 | $0.04 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $158.40 | $17.60 |
| Where the session’s cost goes | ||
| Cache writes | $0.56 | $0.06 |
| Cache reads | $0.52 | $0.06 |
| Uncached input | $0.14 | $0.02 |
| Output | $0.22 | $0.03 |
- caching saves on the session with GLM-5.3 (61%)
- $2.28
- caching saves on the session with GLM-5.3-Flash (60%)
- $0.24
How much cheaper is GLM-5.3-Flash than GLM-5.3?
At Z.ai's own rates, GLM-5.3-Flash charges $0.15 per million input tokens, $0.03 per million cache hits, and $0.50 per million output tokens. GLM-5.3 charges $1.40, $0.26, and $4.40. Every rate is between 8.7x and 9.3x higher on the flagship.
The example workloads follow suit. The agentic coding session costs $1.44 on GLM-5.3 and $0.16 on GLM-5.3-Flash, 9x apart. The large one-off review costs $0.25 against $0.03, and the output-heavy generation $0.39 against $0.04. At 110 sessions a month, the session totals $158.40 against $17.60, a difference of $140.80.
Flash's figures are small enough that rounding to the cent shows. Its session lines, $0.06 of cache writes, $0.06 of cache reads, $0.02 of fresh input, and $0.03 of output, are each rounded, so ratios on a single Flash workload are approximate. The monthly totals are the steadier comparison.
Same limits and caching rules at two price points
The limits match. Both models take 1M tokens of context and write up to 128K tokens of output, both are offered through OpenRouter and in OpenCode, and both use Z.ai's automatic caching, which needs no configuration and carries no fee for writing the cache.
The cache discount is nearly the same share of input: a hit is 18.6% of input on GLM-5.3 and 20% on GLM-5.3-Flash. Caching therefore saves a similar share on each, 61% and 60% of the uncached session, which in dollars is $2.28 on GLM-5.3 and $0.24 on Flash. Z.ai says cached input storage is free for a limited time.
Because the rules match, moving work between the two changes what a session costs, not how it should be structured. The cache example on our homepage shows the session both sets of figures come from.
What Z.ai says each GLM model is for
Z.ai calls GLM-5.3 its "latest flagship model, delivering comprehensive advancements in complex software engineering and agent capabilities," and positions it for long-horizon agent work. Reasoning is always on, at low, high, or max, and Z.ai serves it on OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages endpoints.
GLM-5.3-Flash is Z.ai's low-cost, natively multimodal model, released on August 26, 2026, twelve days after GLM-5.3. Z.ai says it outperforms GLM-5.2 at a tenth of the price, and describes native multimodal visual coding, in which the model looks at interfaces and rendered results to test and improve its work. Z.ai also credits a hybrid of sparse and linear attention with cutting attention compute and KV cache size.
Flash already has wide use. By OpenRouter's count, summed across nine languages, Flash was the most-used model for programming over the week before September 28, 2026. That is a usage figure rather than a quality claim, but it shows how many developers already point code at it.
Licenses, self-hosting, and tracking both
Both models ship open weights, under different terms. GLM-5.3-Flash uses the MIT license. GLM-5.3 uses Z.ai's own GLM-5.3 license, so read its terms before building on the weights. This page does not estimate hosting costs for either.
The prices here are Z.ai's own API rates. OpenRouter sends GLM requests to one of several providers, whose prices can differ. EveryToken prices both GLM models from OpenRouter's catalog when you use them through OpenRouter, so a mix of flagship and Flash requests shows up side by side; it does not price calls made directly to Z.ai's API.
Prompt caching
How Z.ai bills cached tokens
Z.ai
Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.
A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.
Source: Z.ai docs: Context caching
Your own numbers
See what GLM-5.3 and GLM-5.3-Flash really cost you.
everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Does GLM-5.3 really cost nine times as much as GLM-5.3-Flash?
Close to it. Every list rate is 8.7x to 9.3x higher on GLM-5.3, and the example agentic session costs $1.44 against $0.16, a 9x gap. At 110 sessions a month that is $158.40 against $17.60.
Do GLM-5.3 and GLM-5.3-Flash have the same limits?
Yes. GLM-5.3 and GLM-5.3-Flash both take 1M tokens of context and write up to 128K tokens of output. Price, license, and the work Z.ai aims each at are what separate them.
Which license does each GLM model use?
GLM-5.3-Flash is released under the MIT license. GLM-5.3 uses Z.ai's own GLM-5.3 license. Both GLM models have open weights, so either can be self-hosted within its license terms.
Do both models cache the same way?
Yes. Z.ai caches repeated context automatically for both and lists no fee for writing the cache. A hit costs $0.26 per million on GLM-5.3 and $0.03 on GLM-5.3-Flash.