Skip to content

Model comparison

Qwen3.8-Max vs GLM-5.3: list prices, caching, and weights

GLM-5.3 costs $1.44 per cached coding session against $1.80 on Qwen3.8-Max, even though cache hits cost about the same. Rates, caching, weights, and tools.

· Prices as of September 28, 2026

  • Qwen3.8-Max

    Alibaba Qwen · Released August 2, 2026

    Qwen's most capable model, built for long autonomous coding and professional work.

    Qwen3.8-Max facts and comparisons
  • GLM-5.3

    Z.ai · Released August 14, 2026

    Z.ai's flagship for complex software engineering and long-horizon agent work, with open weights.

    GLM-5.3 facts and comparisons

The short answer

GLM-5.3 costs less than Qwen3.8-Max on every example workload, $1.44 against $1.80 for the agentic coding session and $0.25 against $0.36 for a large one-off review, because its input and output rates are $1.40 and $4.40 per million against $2 and $6. Cache hits are nearly identical, $0.26 on GLM-5.3 and $0.25 on Qwen3.8-Max, so the session gap, 20%, is narrower than the rate gap. GLM-5.3 ships open weights under Z.ai's own license, while Qwen3.8-Max is a closed API model whose base variant has open weights.

Choose Qwen3.8-Max if

  • You like Qwen's account of training it with reinforcement learning across harnesses including Claude Code and Codex.
  • You want explicit caching as an option: a cache_control breakpoint on Qwen3.8-Max reads at $0.17 per million.
  • You want 131K of output per request, slightly more than GLM-5.3's 128K.

Choose GLM-5.3 if

  • You want lower list prices: $1.40 input and $4.40 output per million tokens.
  • You want the flagship model's own weights, which Z.ai publishes under its GLM-5.3 license.
  • You want one model reachable through OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages endpoints.
  • You prefer a subscription, since Z.ai also sells GLM-5.3 through its GLM Coding Plan.

Side by side

Specs and prices

FactQwen3.8-MaxGLM-5.3
MakerAlibaba QwenZ.ai
API model idqwen3.8-maxglm-5.3
ReleasedAugust 2, 2026August 14, 2026
StatusCurrentCurrent
Context window1M tokens1M tokens
Max output131K tokens128K tokens
Open weightsNoYes
Input, per 1M tokens$2$1.40
Cache hit, per 1M$0.25$0.26
Cache write, per 1M$2 (same as input)$1.40 (same as input)
Output, per 1M tokens$6$4.40
Runs inOpenCode and OpenRouterOpenCode and OpenRouter

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Qwen3.8-Max: Qwen Cloud prices. Alibaba Cloud Model Studio's international scope lists the same input and output rates.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadQwen3.8-MaxGLM-5.3
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$1.80$1.44
Large one-off review, 150K input with no cache hits, 10K output$0.36$0.25
Output-heavy generation, 30K input, 80K output$0.54$0.39
A month of sessions, 110 sessions: 5 a day, 22 working days$198.00$158.40
Where the session’s cost goes
Cache writes$0.80$0.56
Cache reads$0.50$0.52
Uncached input$0.20$0.14
Output$0.30$0.22
caching saves on the session with Qwen3.8-Max (66%)
$3.50
caching saves on the session with GLM-5.3 (61%)
$2.28

Where GLM-5.3's 20% saving comes from

GLM-5.3 lists $1.40 per million input tokens and $4.40 per million output tokens. Qwen3.8-Max lists $2 and $6. Input is 30% cheaper on GLM-5.3 and output 27% cheaper, close to what the uncached workloads show: $0.25 against $0.36 for the large one-off review and $0.39 against $0.54 for the output-heavy generation.

Cache hits don't follow. Z.ai charges $0.26 per million for a GLM-5.3 hit, 18.6% of its input price, and Qwen charges $0.25 for a Qwen3.8-Max hit, 12.5% of input. In the example session the 2M cached tokens cost $0.52 on GLM-5.3 and $0.50 on Qwen3.8-Max, the one line where Qwen3.8-Max comes out ahead.

Both bill cache writes as ordinary input, so writes track the input gap: $0.56 on GLM-5.3 against $0.80. The session totals $1.44 against $1.80, 20% less on GLM-5.3, and a month of 110 sessions $158.40 against $198.00. Because Qwen3.8-Max's hit is a smaller fraction of its input, caching saves it more, 66% of its uncached session cost against 61% on GLM-5.3.

Implicit and explicit caching on Qwen3.8-Max

Qwen's implicit caching runs automatically and can't be switched off, and written tokens cost ordinary input. That is how the tables price it. Qwen also offers explicit caching, marked with cache_control, which lasts 5 minutes and costs $2.50 per million tokens to create and $0.17 per million to read.

Explicit caching trades a write premium for cheaper reads. At $0.17, an explicit hit costs less than GLM-5.3's $0.26, so a Qwen3.8-Max session that rereads one fixed prefix many times while the explicit cache lives could narrow the gap. The tables leave that out, because whether it pays depends on how often your tool reuses the same prefix.

On the other side, Z.ai caches repeated context automatically, with no configuration and no write fee, and says cached input storage is free for a limited time.

Weights, context, and the tools that carry them

Z.ai publishes GLM-5.3's weights under its own GLM-5.3 license. Qwen3.8-Max is a closed API model, and the open weights Qwen publishes are for its base variant, Qwen3.8-2.4T-A95B. The Qwen3.8-Max prices here are Qwen Cloud's, and Alibaba Cloud Model Studio's international scope lists the same input and output rates.

Both accept 1M tokens of context, and Qwen3.8-Max writes up to 131K per request against 128K on GLM-5.3. Neither price list shows a long-context tier. GLM-5.3 always reasons, at low, high, or max, and the level you pick changes how much output a task produces.

Both run in OpenCode and OpenRouter, and neither is listed in Cursor or GitHub Copilot. On OpenRouter, an open-weight model like GLM-5.3 is served by one of several providers, whose prices can differ from Z.ai's own API, and the tables use each maker's own price.

How Qwen and Z.ai position their flagships

Qwen calls Qwen3.8-Max "the most capable model in the Qwen family to date" and built it for long autonomous coding and professional work. It describes a self-evolving harness built during an autonomous coding run of more than 10 days, and reinforcement learning across harnesses including Claude Code and Codex. Alibaba Cloud Model Studio names it the pick for the strongest reasoning in coding tools.

Z.ai describes GLM-5.3 as "Z.ai's latest flagship model, delivering comprehensive advancements in complex software engineering and agent capabilities." It serves the model through OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages endpoints, and also sells it with the GLM Coding Plan subscription.

EveryToken covers Qwen3.8-Max and GLM-5.3 only when they run through OpenRouter, at OpenRouter's catalog prices, and not through Qwen Cloud or Z.ai's own API.

Prompt caching

How each maker bills cached tokens

Alibaba Qwen

Implicit caching is automatic and can't be turned off. A hit on Qwen3.8-Max costs $0.25 per million tokens, and written tokens cost ordinary input.

Explicit caching, marked with cache_control, lasts 5 minutes. On Qwen3.8-Max it costs $2.50 per million tokens to create and $0.17 per million to read. This blog prices Qwen with implicit caching.

Source: Alibaba Cloud Model Studio: Context cache

Z.ai

Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.

A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.

Source: Z.ai docs: Context caching

Your own numbers

See what Qwen3.8-Max and GLM-5.3 really cost you.

everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is GLM-5.3 cheaper than Qwen3.8-Max?

Yes, on every example workload. Input and output cost $1.40 and $4.40 per million on GLM-5.3 against $2 and $6, and the agentic coding session costs $1.44 against $1.80.

Are GLM-5.3 and Qwen3.8-Max open-weight models?

GLM-5.3 is, under Z.ai's own GLM-5.3 license. Qwen3.8-Max is closed, though Qwen publishes open weights for its base variant, Qwen3.8-2.4T-A95B.

How does prompt caching differ between the two?

Both cache automatically and bill written tokens as input. A hit costs $0.25 per million on Qwen3.8-Max and $0.26 on GLM-5.3. Qwen also offers explicit caching at $2.50 per million to create and $0.17 to read, lasting 5 minutes.

Which tools offer both models?

OpenCode and OpenRouter. Neither is listed in Cursor or GitHub Copilot, but GLM-5.3 exposes Anthropic Messages and OpenAI Responses endpoints alongside OpenAI Chat Completions.

  • GLM-5.3 vs GLM-5.3-Flash

    GLM-5.3-Flash costs about a ninth of GLM-5.3 at Z.ai's rates: $0.16 against $1.44 for an agentic coding session. Same limits, same caching, different jobs.

  • DeepSeek-V4-Pro vs GLM-5.3

    DeepSeek-V4-Pro and GLM-5.3 list input and output within 11% of each other, but a DeepSeek cache hit costs $0.044 against $0.26. What that does to a session.

  • GLM-5.3 vs Claude Sonnet 5

    GLM-5.3 undercuts Claude Sonnet 5 by 40% on an agentic coding session, $1.44 against $2.40. Most of it comes from Anthropic's cache-write premium.

  • GLM-5.3 vs Claude Opus 5.5

    GLM-5.3 costs about a third of Claude Opus 5.5 on an agentic coding session, $1.44 against $4.40, yet its cache hits cost more. Where the gap comes from.

  • GLM-5.3 vs Gemini 3.1 Pro Preview

    Neither Z.ai nor Google charges extra to write the cache, so GLM-5.3 and Gemini 3.1 Pro Preview split on output: $4.40 against $12 per million tokens.

  • GLM-5.3 vs GPT-6 Sol

    GLM-5.3 costs 31% less than GPT-6 Sol on an agentic coding session, $1.44 against $2.10. How OpenAI's write fee and 272K rule compare with Z.ai's rates.