Z.ai
GLM-5.3: price, context window, and caching
Z.ai's flagship for complex software engineering and long-horizon agent work, with open weights.
Released August 14, 2026 · Prices as of September 28, 2026
In Z.ai’s words
“GLM-5.3 is Z.ai's latest flagship model, delivering comprehensive advancements in complex software engineering and agent capabilities.”
Facts
Specs and prices
| Fact | GLM-5.3 |
|---|---|
| Maker | Z.ai |
| API model id | glm-5.3 |
| Released | August 14, 2026 |
| Status | Current |
| Context window | 1M tokens |
| Max output | 128K tokens |
| Open weights | Yes |
| Input, per 1M tokens | $1.40 |
| Cache hit, per 1M | $0.26 |
| Cache write, per 1M | $1.40 (same as input) |
| Output, per 1M tokens | $4.40 |
| Runs in | OpenCode and OpenRouter |
Standard API rates in US dollars, as published by the maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included.
Good to know
- Open weights under Z.ai's own GLM-5.3 license.
- Reasoning is always on, at low, high, or max.
- Also sold with Z.ai's GLM Coding Plan subscription.
Cost
What typical work costs
Example token counts at GLM-5.3’s published rates. On the agentic session, caching saves $2.28 against billing every token as ordinary input.
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $1.44 |
|---|---|
| Large one-off review, 150K input with no cache hits, 10K output | $0.25 |
| Output-heavy generation, 30K input, 80K output | $0.39 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $158.40 |
Prompt caching
How Z.ai bills cached tokens
Z.ai
Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.
A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.
Source: Z.ai docs: Context caching
GLM-5.3 compared
DeepSeek-V4-Pro vs GLM-5.3
DeepSeek-V4-Pro and GLM-5.3 list input and output within 11% of each other, but a DeepSeek cache hit costs $0.044 against $0.26. What that does to a session.
GLM-5.3 vs Claude Sonnet 5
GLM-5.3 undercuts Claude Sonnet 5 by 40% on an agentic coding session, $1.44 against $2.40. Most of it comes from Anthropic's cache-write premium.
GLM-5.3 vs Claude Opus 5.5
GLM-5.3 costs about a third of Claude Opus 5.5 on an agentic coding session, $1.44 against $4.40, yet its cache hits cost more. Where the gap comes from.
GLM-5.3 vs Gemini 3.1 Pro Preview
Neither Z.ai nor Google charges extra to write the cache, so GLM-5.3 and Gemini 3.1 Pro Preview split on output: $4.40 against $12 per million tokens.
GLM-5.3 vs GLM-5.3-Flash
GLM-5.3-Flash costs about a ninth of GLM-5.3 at Z.ai's rates: $0.16 against $1.44 for an agentic coding session. Same limits, same caching, different jobs.
GLM-5.3 vs GPT-6 Sol
GLM-5.3 costs 31% less than GPT-6 Sol on an agentic coding session, $1.44 against $2.10. How OpenAI's write fee and 272K rule compare with Z.ai's rates.
Kimi K3 vs GLM-5.3
Kimi K3 and GLM-5.3 are both open-weight flagships. A coding session costs $2.85 on Kimi K3 and $1.44 on GLM-5.3, yet their cache hits nearly match.
Qwen3.8-Max vs GLM-5.3
GLM-5.3 costs $1.44 per cached coding session against $1.80 on Qwen3.8-Max, even though cache hits cost about the same. Rates, caching, weights, and tools.
Your own numbers
See what GLM-5.3 really costs you.
everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.