Cost guide
What GLM-5.3 costs for coding: API, cache, and plan
What a coding session costs on GLM-5.3 and GLM-5.3-Flash at Z.ai's API rates, how Z.ai's caching changes it, and when the Coding Plan is cheaper.
· Facts checked October 3, 2026
- Claude Code
- OpenCode
- OpenRouter
The short answer
At Z.ai's API rates, this site's standard agentic coding session costs $1.44 on GLM-5.3 and $0.16 on GLM-5.3-Flash, against $2.40 on Claude Sonnet 5. Z.ai charges nothing extra to write the cache, but a cache hit still costs about 19% of the input price, and GLM-5.3 reasons at its maximum effort by default, which adds output.
At a glance
| GLM-5.3 rates | $1.40 input, $0.26 cached input, $4.40 output per million tokens. |
|---|---|
| GLM-5.3-Flash rates | $0.15 input, $0.03 cached input, $0.50 output per million tokens. |
| Agentic session | $1.44 on GLM-5.3, $0.16 on GLM-5.3-Flash. |
| Cache writes | No write fee; cached storage is free for a limited time, with no end date given. |
| Reasoning default | Max effort. Z.ai cites about 75K output tokens per task at max against about 50K at high. |
Three workloads, priced
The session workload used across this blog is 100,000 tokens of fresh input, 400,000 written to the cache, 2 million read from it, and 50,000 of output. On GLM-5.3 that comes to $1.44: $0.14 of input, $0.56 for the tokens written to cache at the ordinary input price, $0.52 of cache reads, and $0.22 of output. Without caching, the same tokens would cost $3.72.
A large one-off review, 150,000 input tokens with no cache hits and 10,000 output, costs $0.25. An output-heavy generation, 30,000 in and 80,000 out, costs $0.39. On GLM-5.3-Flash the three workloads cost $0.16, $0.03, and $0.04, which is why the GLM-5.3 vs GLM-5.3-Flash comparison is mostly a question of quality.
How Z.ai's caching differs
Anthropic charges a premium to write the cache and then bills hits at a tenth of input. Z.ai does the opposite: writes cost ordinary input, and a hit on GLM-5.3 costs $0.26 per million, about 19% of the $1.40 input price. Caching therefore never costs more than not caching on GLM, but each read saves less than on Claude, so the share of a GLM session spent on cache reads is larger. The cache write premium entry explains the Claude side.
Caching is automatic, with no configuration, and Z.ai lists cached storage as free for a limited time without an end date. On a Coding Plan, Z.ai keeps reasoning from earlier turns by default, which raises cache hits.
Reasoning effort is a cost setting
GLM-5.3 always reasons, at low, high, or max, and max is the default. Z.ai's own coding evaluation cites about 75,000 output tokens per task at max against about 50,000 at high. At $4.40 per million output tokens, those extra 25,000 tokens add about $0.11 per task, so routine edits at high effort cost noticeably less. In Claude Code, /effort sets the level, as the Claude Code setup guide explains.
API or Coding Plan
A hundred and ten sessions a month, five a day over twenty-two working days, would cost $158.40 on GLM-5.3 at API rates. The GLM Coding Plan meters the same session at 805 credits at peak hours, so Lite's weekly allowance covers about twelve of them for $18 a month, Pro about seventy-four for $80. For steady use the plan is far cheaper per session; for occasional use, or to mix models, paying per token through Z.ai or OpenRouter avoids a monthly fee.
Your own numbers
See what your AI coding really costs.
everyaitoken reads Claude Code, Codex, Cursor, OpenRouter, Gemini CLI, and OpenCode on your Mac. It shows limits with reset times, API-equivalent cost, and what caching saved or cost, in the menu bar and 17 widgets. $9 once.
FAQ
Questions
How much does GLM-5.3 cost per million tokens?
On Z.ai's API, $1.40 per million input tokens, $0.26 per million cached input tokens, and $4.40 per million output tokens. OpenRouter lists the same base price, but individual providers vary.
Does Z.ai charge for prompt cache writes?
No. Written tokens cost the ordinary input price, cache hits cost $0.26 per million on GLM-5.3, and cached storage is free for a limited time.
Is GLM-5.3 cheaper than Claude Sonnet 5 for coding?
At list prices, yes: the standard session costs $1.44 on GLM-5.3 and $2.40 on Claude Sonnet 5. The GLM-5.3 vs Claude Sonnet 5 comparison breaks down where the gap comes from.
Do reasoning tokens make GLM-5.3 more expensive?
They add output. Z.ai cites about 75,000 output tokens per task at max effort and about 50,000 at high, so choosing high for routine work lowers cost.
Sources
Checked October 3, 2026. Plans and limits change; the linked pages are the authority.