Skip to content

Z.ai

GLM-5.3-Flash: price, context window, and caching

Z.ai's low-cost, natively multimodal model, which Z.ai says outperforms GLM-5.2 at a tenth of the price.

Released August 26, 2026 · Prices as of September 28, 2026

In Z.ai’s words

“it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price”

Z.ai: GLM-5.3-Flash

What Z.ai says it’s good at

  • Native multimodal visual coding: it looks at interfaces and rendered results to test and improve its work Source
  • Hybrid sparse and linear attention that cuts attention compute and KV cache size Source

Facts

Specs and prices

FactGLM-5.3-Flash
MakerZ.ai
API model idglm-5.3-flash
ReleasedAugust 26, 2026
StatusCurrent
Context window1M tokens
Max output128K tokens
Open weightsYes
Input, per 1M tokens$0.15
Cache hit, per 1M$0.03
Cache write, per 1M$0.15 (same as input)
Output, per 1M tokens$0.50
Runs inOpenCode and OpenRouter

Standard API rates in US dollars, as published by the maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included.

Good to know

  • Open weights under the MIT license.
  • The most-used model for programming on OpenRouter over the week before September 28, 2026, summed across nine languages.

Cost

What typical work costs

Example token counts at GLM-5.3-Flash’s published rates. On the agentic session, caching saves $0.24 against billing every token as ordinary input.

Example workload costs for GLM-5.3-Flash
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$0.16
Large one-off review, 150K input with no cache hits, 10K output$0.03
Output-heavy generation, 30K input, 80K output$0.04
A month of sessions, 110 sessions: 5 a day, 22 working days$17.60

Prompt caching

How Z.ai bills cached tokens

Z.ai

Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.

A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.

Source: Z.ai docs: Context caching

GLM-5.3-Flash compared

  • GLM-5.3-Flash vs Claude Sonnet 5

    GLM-5.3-Flash costs $0.16 for a cached coding session that costs $2.40 on Claude Sonnet 5. What drives a 15x gap, and what price alone can't tell you.

  • DeepSeek-V4.1-Flash vs GLM-5.3-Flash

    GLM-5.3-Flash lists half DeepSeek-V4.1-Flash's input price, but GLM's cache hits cost 5x as much. A cached coding session: $0.16 against $0.22 at peak.

  • GLM-5.3 vs GLM-5.3-Flash

    GLM-5.3-Flash costs about a ninth of GLM-5.3 at Z.ai's rates: $0.16 against $1.44 for an agentic coding session. Same limits, same caching, different jobs.

  • GLM-5.3-Flash vs Claude Haiku 4.5

    GLM-5.3-Flash charges $0.50 per million output tokens against $5 on Claude Haiku 4.5, and $0.16 against $1.20 for a cached coding session. Where each fits.

  • GLM-5.3-Flash vs Gemini 3.8 Flash

    GLM-5.3-Flash costs $0.16 for a cached coding session against $0.71 on Gemini 3.8 Flash, whose introductory rates end December 31, 2026. What else differs.

  • GLM-5.3-Flash vs GPT-6 Luna

    GPT-6 Luna and GLM-5.3-Flash share a $0.50 output rate, yet a cached coding session costs $0.11 on Luna and $0.16 on GLM-5.3-Flash. Cache hits explain why.

  • MiniMax M3 vs GLM-5.3-Flash

    GLM-5.3-Flash costs about half as much as MiniMax M3 on every line, $0.16 against $0.33 per cached coding session. Output limits and licenses differ too.

Your own numbers

See what GLM-5.3-Flash really costs you.

everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math