Model comparison
GLM-5.3 or Claude Sonnet 5: where a 40% saving comes from
GLM-5.3 undercuts Claude Sonnet 5 by 40% on an agentic coding session, $1.44 against $2.40. Most of it comes from Anthropic's cache-write premium.
· Prices as of September 28, 2026
GLM-5.3
Z.ai · Released August 14, 2026
Z.ai's flagship for complex software engineering and long-horizon agent work, with open weights.
GLM-5.3 facts and comparisonsClaude Sonnet 5
Anthropic · Released June 30, 2026
Anthropic's balance of speed and intelligence, and a drop-in upgrade for Claude Sonnet 4.6. Anthropic pitches it as close to Claude Opus 4.8 at a lower price.
Claude Sonnet 5 facts and comparisons
The short answer
GLM-5.3 costs 40% less than Claude Sonnet 5 on the example agentic coding session, $1.44 against $2.40, and 55% less on output-heavy work, although Sonnet 5 charges less for each cache hit. Choose Sonnet 5 if you want it in Claude Code, Cursor, or GitHub Copilot; choose GLM-5.3 for the lower rates and open weights, reached through OpenRouter or OpenCode.
Choose GLM-5.3 if
- You want output at $4.40 per million tokens rather than $10, which counts most when a model writes long files or large diffs.
- Your sessions write a lot of new context to the cache, which Z.ai bills at the $1.40 input rate with no premium.
- You want open weights under Z.ai's GLM-5.3 license, or a GLM Coding Plan subscription instead of per-token billing.
Choose Claude Sonnet 5 if
- Claude Code is your main tool: its sonnet alias points to Claude Sonnet 5 on the Anthropic API.
- You want the same model across Claude Code, Cursor, and GitHub Copilot, all of which offer Sonnet 5.
- Your context is stable and rereads dominate, where a Sonnet 5 cache hit costs $0.20 per million against $0.26 on GLM-5.3, though its writes and output still cost more.
- You are moving off Claude Sonnet 4.6, which Anthropic says Sonnet 5 replaces as a drop-in upgrade.
Side by side
Specs and prices
| Fact | GLM-5.3 | Claude Sonnet 5 |
|---|---|---|
| Maker | Z.ai | Anthropic |
| API model id | glm-5.3 | claude-sonnet-5 |
| Released | August 14, 2026 | June 30, 2026 |
| Status | Current | Current |
| Context window | 1M tokens | 1M tokens |
| Max output | 128K tokens | 128K tokens |
| Open weights | Yes | No |
| Input, per 1M tokens | $1.40 | $2 |
| Cache hit, per 1M | $0.26 | $0.20 |
| Cache write, per 1M | $1.40 (same as input) | $2.50 (5-minute), $4 (1-hour) |
| Output, per 1M tokens | $4.40 | $10 |
| Runs in | OpenCode and OpenRouter | Claude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker (GLM-5.3: September 28, 2026; Claude Sonnet 5: September 26, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Claude Sonnet 5: The full 1M context window is billed at standard rates. The launch price became the standard price on August 10, 2026, and a planned increase was cancelled.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | GLM-5.3 | Claude Sonnet 5 |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $1.44 | $2.40 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.25 | $0.40 |
| Output-heavy generation, 30K input, 80K output | $0.39 | $0.86 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $158.40 | $264.00 |
| Where the session’s cost goes | ||
| Cache writes | $0.56 | $1.30 |
| Cache reads | $0.52 | $0.40 |
| Uncached input | $0.14 | $0.20 |
| Output | $0.22 | $0.50 |
- caching saves on the session with GLM-5.3 (61%)
- $2.28
- caching saves on the session with Claude Sonnet 5 (56%)
- $3.10
Where the $0.96 difference per session comes from
Claude Sonnet 5 lists $2 per million input tokens and $10 per million output tokens. GLM-5.3 lists $1.40 and $4.40. On input the gap is modest, 30% less on GLM-5.3, while output is 56% cheaper. Those two ratios explain most of what follows.
Split the example session into its parts and the cache writes carry the gap. The session writes 400K tokens to the cache: $0.56 on GLM-5.3, which bills writes as ordinary input, and $1.30 on Sonnet 5, where Anthropic charges 1.25x input for a 5-minute write and 2x for a 1-hour write. That is $0.74 of the $0.96 difference. Output adds another $0.28 and fresh input $0.06.
Cache reads push back by $0.12. The session reads 2M tokens from the cache, which costs $0.52 on GLM-5.3 and $0.40 on Sonnet 5. Net of all four lines, GLM-5.3 comes to $1.44 and Sonnet 5 to $2.40, or $158.40 against $264.00 over 110 sessions a month.
Which model charges less for a cache hit?
Sonnet 5 does. Anthropic prices a hit at 0.1x input, $0.20 per million. Z.ai's hit on GLM-5.3 is $0.26, which is 18.6% of its input price, a shallower discount than the 90% off that Sonnet 5 gives. On the rate that covers most tokens in a long agentic session, the model with the higher list price is 23% cheaper.
That is not enough to flip the result, because Sonnet 5's write premium costs more than its read discount saves. The balance depends on setup, though. In Claude Code the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key, and a Sonnet 5 session written only at the 5-minute rate would move closer to GLM-5.3. By Anthropic's own rule of thumb, a Sonnet 5 cache write pays for itself after one read at the 5-minute lifetime and after two at 1 hour.
Against sending every token uncached, caching saves $2.28 on GLM-5.3, or 61%, and $3.10 on Sonnet 5, or 56%. For GLM-5.3, Z.ai caches repeated context automatically, with no configuration, and says cached input storage is free for a limited time. The cache example on our homepage walks through the same session.
Tools, open weights, and reasoning defaults
You can pick Sonnet 5 in Claude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot, while GLM-5.3 is offered through OpenRouter and OpenCode. Z.ai's API also accepts GLM-5.3 requests in the OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages formats. Claude Code is built around Anthropic's models, and pointing it at another provider takes custom configuration that this page does not cover.
GLM-5.3 has open weights under Z.ai's own GLM-5.3 license, so self-hosting is possible, at a hardware cost this page does not try to estimate. The rates here are Z.ai's own API prices. When GLM-5.3 runs through OpenRouter, the request goes to one of several providers, and their prices can differ from Z.ai's.
Defaults shape output length, and output length shapes cost. Sonnet 5 has adaptive thinking on by default at high effort. GLM-5.3 always reasons, at low, high, or max. Anthropic's newer tokenizer counts about 1.0 to 1.35x as many tokens as Claude Sonnet 4.6 for the same text, and Z.ai's tokenizer differs again, so the same prompt will not produce the same token count on both.
How Z.ai and Anthropic describe the two models
Anthropic says "Claude Sonnet 5 is built to be the most agentic Sonnet model yet" and pitches it as close to Claude Opus 4.8 at lower prices. For Sonnet 5 it names planning and using tools like browsers and terminals on its own as strengths. Sonnet 5's price has settled: the launch price became the standard price on August 10, 2026, and a planned increase was cancelled.
Z.ai calls GLM-5.3 its latest flagship, with "comprehensive advancements in complex software engineering and agent capabilities." Released on August 14, 2026, it is about six weeks newer than Sonnet 5. EveryToken prices Sonnet 5 from your Claude Code, Cursor, and OpenCode history, and GLM-5.3 when you reach it through OpenRouter, so the two can be compared on your own work rather than on an example.
Prompt caching
How each maker bills cached tokens
Z.ai
Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.
A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.
Source: Z.ai docs: Context caching
Anthropic
Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.
A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.
Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.
In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.
Source: Anthropic: Prompt caching
Your own numbers
See what GLM-5.3 and Claude Sonnet 5 really cost you.
everyaitoken reads your OpenRouter, Claude Code, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is GLM-5.3 cheaper than Claude Sonnet 5?
Yes on input, output, and cache writes, and on all three example workloads: $1.44 against $2.40 for the agentic session, $0.25 against $0.40 for the uncached review, and $0.39 against $0.86 for output-heavy generation. Sonnet 5 costs less per cache hit, $0.20 against $0.26 per million.
Why is the session gap smaller than the output gap?
Output is 2.3x dearer on Sonnet 5, but most of the session's tokens are cache reads, where Sonnet 5 is cheaper. The generation workload is mostly output, so it shows a 2.2x gap, while the session shows 1.7x.
Is GLM-5.3 available in Cursor or GitHub Copilot?
No. Neither offers GLM-5.3, while both offer Claude Sonnet 5. GLM-5.3 is available through OpenRouter and in OpenCode.
Do GLM-5.3 and Claude Sonnet 5 have the same context window?
Yes. GLM-5.3 and Sonnet 5 both accept 1M tokens of context and write up to 128K tokens of output, so limits do not separate them. The differences are in price, caching, and where each model runs.
Sources
- Z.ai docs: Pricing
- Z.ai docs: GLM-5.3
- Z.ai: GLM-5.3
- OpenRouter: GLM-5.3
- OpenCode docs: Zen
- Anthropic: Pricing
- Anthropic docs: Claude Sonnet 5
- Anthropic: Introducing Claude Sonnet 5
- Claude Code docs: Model configuration
- Cursor docs: Claude Sonnet 5
- OpenRouter: Claude Sonnet 5
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Z.ai docs: Context caching
- Anthropic: Prompt caching