Model comparison
Qwen3.8-Max vs GLM-5.3: list prices, caching, and weights
GLM-5.3 costs $1.44 per cached coding session against $1.80 on Qwen3.8-Max, even though cache hits cost about the same. Rates, caching, weights, and tools.
· Prices as of September 28, 2026
Qwen3.8-Max
Alibaba Qwen · Released August 2, 2026
Qwen's most capable model, built for long autonomous coding and professional work.
Qwen3.8-Max facts and comparisonsGLM-5.3
Z.ai · Released August 14, 2026
Z.ai's flagship for complex software engineering and long-horizon agent work, with open weights.
GLM-5.3 facts and comparisons
The short answer
GLM-5.3 costs less than Qwen3.8-Max on every example workload, $1.44 against $1.80 for the agentic coding session and $0.25 against $0.36 for a large one-off review, because its input and output rates are $1.40 and $4.40 per million against $2 and $6. Cache hits are nearly identical, $0.26 on GLM-5.3 and $0.25 on Qwen3.8-Max, so the session gap, 20%, is narrower than the rate gap. GLM-5.3 ships open weights under Z.ai's own license, while Qwen3.8-Max is a closed API model whose base variant has open weights.
Choose Qwen3.8-Max if
- You like Qwen's account of training it with reinforcement learning across harnesses including Claude Code and Codex.
- You want explicit caching as an option: a cache_control breakpoint on Qwen3.8-Max reads at $0.17 per million.
- You want 131K of output per request, slightly more than GLM-5.3's 128K.
Choose GLM-5.3 if
- You want lower list prices: $1.40 input and $4.40 output per million tokens.
- You want the flagship model's own weights, which Z.ai publishes under its GLM-5.3 license.
- You want one model reachable through OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages endpoints.
- You prefer a subscription, since Z.ai also sells GLM-5.3 through its GLM Coding Plan.
Side by side
Specs and prices
| Fact | Qwen3.8-Max | GLM-5.3 |
|---|---|---|
| Maker | Alibaba Qwen | Z.ai |
| API model id | qwen3.8-max | glm-5.3 |
| Released | August 2, 2026 | August 14, 2026 |
| Status | Current | Current |
| Context window | 1M tokens | 1M tokens |
| Max output | 131K tokens | 128K tokens |
| Open weights | No | Yes |
| Input, per 1M tokens | $2 | $1.40 |
| Cache hit, per 1M | $0.25 | $0.26 |
| Cache write, per 1M | $2 (same as input) | $1.40 (same as input) |
| Output, per 1M tokens | $6 | $4.40 |
| Runs in | OpenCode and OpenRouter | OpenCode and OpenRouter |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Qwen3.8-Max: Qwen Cloud prices. Alibaba Cloud Model Studio's international scope lists the same input and output rates.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Qwen3.8-Max | GLM-5.3 |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $1.80 | $1.44 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.36 | $0.25 |
| Output-heavy generation, 30K input, 80K output | $0.54 | $0.39 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $198.00 | $158.40 |
| Where the session’s cost goes | ||
| Cache writes | $0.80 | $0.56 |
| Cache reads | $0.50 | $0.52 |
| Uncached input | $0.20 | $0.14 |
| Output | $0.30 | $0.22 |
- caching saves on the session with Qwen3.8-Max (66%)
- $3.50
- caching saves on the session with GLM-5.3 (61%)
- $2.28
Where GLM-5.3's 20% saving comes from
GLM-5.3 lists $1.40 per million input tokens and $4.40 per million output tokens. Qwen3.8-Max lists $2 and $6. Input is 30% cheaper on GLM-5.3 and output 27% cheaper, close to what the uncached workloads show: $0.25 against $0.36 for the large one-off review and $0.39 against $0.54 for the output-heavy generation.
Cache hits don't follow. Z.ai charges $0.26 per million for a GLM-5.3 hit, 18.6% of its input price, and Qwen charges $0.25 for a Qwen3.8-Max hit, 12.5% of input. In the example session the 2M cached tokens cost $0.52 on GLM-5.3 and $0.50 on Qwen3.8-Max, the one line where Qwen3.8-Max comes out ahead.
Both bill cache writes as ordinary input, so writes track the input gap: $0.56 on GLM-5.3 against $0.80. The session totals $1.44 against $1.80, 20% less on GLM-5.3, and a month of 110 sessions $158.40 against $198.00. Because Qwen3.8-Max's hit is a smaller fraction of its input, caching saves it more, 66% of its uncached session cost against 61% on GLM-5.3.
Implicit and explicit caching on Qwen3.8-Max
Qwen's implicit caching runs automatically and can't be switched off, and written tokens cost ordinary input. That is how the tables price it. Qwen also offers explicit caching, marked with cache_control, which lasts 5 minutes and costs $2.50 per million tokens to create and $0.17 per million to read.
Explicit caching trades a write premium for cheaper reads. At $0.17, an explicit hit costs less than GLM-5.3's $0.26, so a Qwen3.8-Max session that rereads one fixed prefix many times while the explicit cache lives could narrow the gap. The tables leave that out, because whether it pays depends on how often your tool reuses the same prefix.
On the other side, Z.ai caches repeated context automatically, with no configuration and no write fee, and says cached input storage is free for a limited time.
Weights, context, and the tools that carry them
Z.ai publishes GLM-5.3's weights under its own GLM-5.3 license. Qwen3.8-Max is a closed API model, and the open weights Qwen publishes are for its base variant, Qwen3.8-2.4T-A95B. The Qwen3.8-Max prices here are Qwen Cloud's, and Alibaba Cloud Model Studio's international scope lists the same input and output rates.
Both accept 1M tokens of context, and Qwen3.8-Max writes up to 131K per request against 128K on GLM-5.3. Neither price list shows a long-context tier. GLM-5.3 always reasons, at low, high, or max, and the level you pick changes how much output a task produces.
Both run in OpenCode and OpenRouter, and neither is listed in Cursor or GitHub Copilot. On OpenRouter, an open-weight model like GLM-5.3 is served by one of several providers, whose prices can differ from Z.ai's own API, and the tables use each maker's own price.
How Qwen and Z.ai position their flagships
Qwen calls Qwen3.8-Max "the most capable model in the Qwen family to date" and built it for long autonomous coding and professional work. It describes a self-evolving harness built during an autonomous coding run of more than 10 days, and reinforcement learning across harnesses including Claude Code and Codex. Alibaba Cloud Model Studio names it the pick for the strongest reasoning in coding tools.
Z.ai describes GLM-5.3 as "Z.ai's latest flagship model, delivering comprehensive advancements in complex software engineering and agent capabilities." It serves the model through OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages endpoints, and also sells it with the GLM Coding Plan subscription.
EveryToken covers Qwen3.8-Max and GLM-5.3 only when they run through OpenRouter, at OpenRouter's catalog prices, and not through Qwen Cloud or Z.ai's own API.
Prompt caching
How each maker bills cached tokens
Alibaba Qwen
Implicit caching is automatic and can't be turned off. A hit on Qwen3.8-Max costs $0.25 per million tokens, and written tokens cost ordinary input.
Explicit caching, marked with cache_control, lasts 5 minutes. On Qwen3.8-Max it costs $2.50 per million tokens to create and $0.17 per million to read. This blog prices Qwen with implicit caching.
Z.ai
Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.
A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.
Source: Z.ai docs: Context caching
Your own numbers
See what Qwen3.8-Max and GLM-5.3 really cost you.
everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is GLM-5.3 cheaper than Qwen3.8-Max?
Yes, on every example workload. Input and output cost $1.40 and $4.40 per million on GLM-5.3 against $2 and $6, and the agentic coding session costs $1.44 against $1.80.
Are GLM-5.3 and Qwen3.8-Max open-weight models?
GLM-5.3 is, under Z.ai's own GLM-5.3 license. Qwen3.8-Max is closed, though Qwen publishes open weights for its base variant, Qwen3.8-2.4T-A95B.
How does prompt caching differ between the two?
Both cache automatically and bill written tokens as input. A hit costs $0.25 per million on Qwen3.8-Max and $0.26 on GLM-5.3. Qwen also offers explicit caching at $2.50 per million to create and $0.17 to read, lasting 5 minutes.
Which tools offer both models?
OpenCode and OpenRouter. Neither is listed in Cursor or GitHub Copilot, but GLM-5.3 exposes Anthropic Messages and OpenAI Responses endpoints alongside OpenAI Chat Completions.