Z.ai
GLM-5.3-Flash: price, context window, and caching
Z.ai's low-cost, natively multimodal model, which Z.ai says outperforms GLM-5.2 at a tenth of the price.
Released August 26, 2026 · Prices as of September 28, 2026
In Z.ai’s words
“it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price”
Facts
Specs and prices
| Fact | GLM-5.3-Flash |
|---|---|
| Maker | Z.ai |
| API model id | glm-5.3-flash |
| Released | August 26, 2026 |
| Status | Current |
| Context window | 1M tokens |
| Max output | 128K tokens |
| Open weights | Yes |
| Input, per 1M tokens | $0.15 |
| Cache hit, per 1M | $0.03 |
| Cache write, per 1M | $0.15 (same as input) |
| Output, per 1M tokens | $0.50 |
| Runs in | OpenCode and OpenRouter |
Standard API rates in US dollars, as published by the maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included.
Good to know
- Open weights under the MIT license.
- The most-used model for programming on OpenRouter over the week before September 28, 2026, summed across nine languages.
Cost
What typical work costs
Example token counts at GLM-5.3-Flash’s published rates. On the agentic session, caching saves $0.24 against billing every token as ordinary input.
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.16 |
|---|---|
| Large one-off review, 150K input with no cache hits, 10K output | $0.03 |
| Output-heavy generation, 30K input, 80K output | $0.04 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $17.60 |
Prompt caching
How Z.ai bills cached tokens
Z.ai
Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.
A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.
Source: Z.ai docs: Context caching
GLM-5.3-Flash compared
GLM-5.3-Flash vs Claude Sonnet 5
GLM-5.3-Flash costs $0.16 for a cached coding session that costs $2.40 on Claude Sonnet 5. What drives a 15x gap, and what price alone can't tell you.
DeepSeek-V4.1-Flash vs GLM-5.3-Flash
GLM-5.3-Flash lists half DeepSeek-V4.1-Flash's input price, but GLM's cache hits cost 5x as much. A cached coding session: $0.16 against $0.22 at peak.
GLM-5.3 vs GLM-5.3-Flash
GLM-5.3-Flash costs about a ninth of GLM-5.3 at Z.ai's rates: $0.16 against $1.44 for an agentic coding session. Same limits, same caching, different jobs.
GLM-5.3-Flash vs Claude Haiku 4.5
GLM-5.3-Flash charges $0.50 per million output tokens against $5 on Claude Haiku 4.5, and $0.16 against $1.20 for a cached coding session. Where each fits.
GLM-5.3-Flash vs Gemini 3.8 Flash
GLM-5.3-Flash costs $0.16 for a cached coding session against $0.71 on Gemini 3.8 Flash, whose introductory rates end December 31, 2026. What else differs.
GLM-5.3-Flash vs GPT-6 Luna
GPT-6 Luna and GLM-5.3-Flash share a $0.50 output rate, yet a cached coding session costs $0.11 on Luna and $0.16 on GLM-5.3-Flash. Cache hits explain why.
MiniMax M3 vs GLM-5.3-Flash
GLM-5.3-Flash costs about half as much as MiniMax M3 on every line, $0.16 against $0.33 per cached coding session. Output limits and licenses differ too.
Your own numbers
See what GLM-5.3-Flash really costs you.
everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.