Skip to content

Z.ai

GLM-5.3: price, context window, and caching

Z.ai's flagship for complex software engineering and long-horizon agent work, with open weights.

Released August 14, 2026 · Prices as of September 28, 2026

In Z.ai’s words

“GLM-5.3 is Z.ai's latest flagship model, delivering comprehensive advancements in complex software engineering and agent capabilities.”

Z.ai docs: GLM-5.3

What Z.ai says it’s good at

  • Complex software engineering and agent capabilities Source
  • Endpoints for OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages Source

Facts

Specs and prices

FactGLM-5.3
MakerZ.ai
API model idglm-5.3
ReleasedAugust 14, 2026
StatusCurrent
Context window1M tokens
Max output128K tokens
Open weightsYes
Input, per 1M tokens$1.40
Cache hit, per 1M$0.26
Cache write, per 1M$1.40 (same as input)
Output, per 1M tokens$4.40
Runs inOpenCode and OpenRouter

Standard API rates in US dollars, as published by the maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included.

Good to know

  • Open weights under Z.ai's own GLM-5.3 license.
  • Reasoning is always on, at low, high, or max.
  • Also sold with Z.ai's GLM Coding Plan subscription.

Cost

What typical work costs

Example token counts at GLM-5.3’s published rates. On the agentic session, caching saves $2.28 against billing every token as ordinary input.

Example workload costs for GLM-5.3
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$1.44
Large one-off review, 150K input with no cache hits, 10K output$0.25
Output-heavy generation, 30K input, 80K output$0.39
A month of sessions, 110 sessions: 5 a day, 22 working days$158.40

Prompt caching

How Z.ai bills cached tokens

Z.ai

Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.

A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.

Source: Z.ai docs: Context caching

GLM-5.3 compared

  • DeepSeek-V4-Pro vs GLM-5.3

    DeepSeek-V4-Pro and GLM-5.3 list input and output within 11% of each other, but a DeepSeek cache hit costs $0.044 against $0.26. What that does to a session.

  • GLM-5.3 vs Claude Sonnet 5

    GLM-5.3 undercuts Claude Sonnet 5 by 40% on an agentic coding session, $1.44 against $2.40. Most of it comes from Anthropic's cache-write premium.

  • GLM-5.3 vs Claude Opus 5.5

    GLM-5.3 costs about a third of Claude Opus 5.5 on an agentic coding session, $1.44 against $4.40, yet its cache hits cost more. Where the gap comes from.

  • GLM-5.3 vs Gemini 3.1 Pro Preview

    Neither Z.ai nor Google charges extra to write the cache, so GLM-5.3 and Gemini 3.1 Pro Preview split on output: $4.40 against $12 per million tokens.

  • GLM-5.3 vs GLM-5.3-Flash

    GLM-5.3-Flash costs about a ninth of GLM-5.3 at Z.ai's rates: $0.16 against $1.44 for an agentic coding session. Same limits, same caching, different jobs.

  • GLM-5.3 vs GPT-6 Sol

    GLM-5.3 costs 31% less than GPT-6 Sol on an agentic coding session, $1.44 against $2.10. How OpenAI's write fee and 272K rule compare with Z.ai's rates.

  • Kimi K3 vs GLM-5.3

    Kimi K3 and GLM-5.3 are both open-weight flagships. A coding session costs $2.85 on Kimi K3 and $1.44 on GLM-5.3, yet their cache hits nearly match.

  • Qwen3.8-Max vs GLM-5.3

    GLM-5.3 costs $1.44 per cached coding session against $1.80 on Qwen3.8-Max, even though cache hits cost about the same. Rates, caching, weights, and tools.

Your own numbers

See what GLM-5.3 really costs you.

everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math