Skip to content

Model comparison

Claude Sonnet 5 or GLM-5.3-Flash? What a 15x gap means

GLM-5.3-Flash costs $0.16 for a cached coding session that costs $2.40 on Claude Sonnet 5. What drives a 15x gap, and what price alone can't tell you.

· Prices as of September 28, 2026

  • GLM-5.3-Flash

    Z.ai · Released August 26, 2026

    Z.ai's low-cost, natively multimodal model, which Z.ai says outperforms GLM-5.2 at a tenth of the price.

    GLM-5.3-Flash facts and comparisons
  • Claude Sonnet 5

    Anthropic · Released June 30, 2026

    Anthropic's balance of speed and intelligence, and a drop-in upgrade for Claude Sonnet 4.6. Anthropic pitches it as close to Claude Opus 4.8 at a lower price.

    Claude Sonnet 5 facts and comparisons

The short answer

GLM-5.3-Flash costs $0.16 for the example agentic coding session against $2.40 on Claude Sonnet 5, a 15x gap that reaches 21.5x on output-heavy work, where Z.ai charges $0.50 per million output tokens against $10. Claude Sonnet 5 is the choice if you work in Claude Code or want the model Anthropic calls its most agentic Sonnet yet. GLM-5.3-Flash suits high-volume or budget-bound work through OpenRouter or OpenCode, and its MIT-licensed weights can be self-hosted.

Choose GLM-5.3-Flash if

  • Cost per session is the constraint: $0.16 against $2.40, or $17.60 against $264.00 over 110 sessions a month.
  • Your workload is output-heavy, and GLM-5.3-Flash bills output at $0.50 per million where Sonnet 5 bills $10.
  • You want the weights as well as the API, under the MIT license.
  • You need Sonnet-sized limits, 1M tokens of context and 128K of output, at a small fraction of the price.

Choose Claude Sonnet 5 if

  • You live in Claude Code and want Anthropic's mid-priced model rather than the Opus default.
  • You want the model Anthropic describes as close to Claude Opus 4.8 at lower prices, built to plan and to drive browsers and terminals on its own.
  • Your team uses Cursor or GitHub Copilot, which offer Sonnet 5 and do not list GLM-5.3-Flash.
  • You are replacing Claude Sonnet 4.6, and Anthropic describes Sonnet 5 as a drop-in upgrade.

Side by side

Specs and prices

FactGLM-5.3-FlashClaude Sonnet 5
MakerZ.aiAnthropic
API model idglm-5.3-flashclaude-sonnet-5
ReleasedAugust 26, 2026June 30, 2026
StatusCurrentCurrent
Context window1M tokens1M tokens
Max output128K tokens128K tokens
Open weightsYesNo
Input, per 1M tokens$0.15$2
Cache hit, per 1M$0.03$0.20
Cache write, per 1M$0.15 (same as input)$2.50 (5-minute), $4 (1-hour)
Output, per 1M tokens$0.50$10
Runs inOpenCode and OpenRouterClaude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker (GLM-5.3-Flash: September 28, 2026; Claude Sonnet 5: September 26, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Claude Sonnet 5: The full 1M context window is billed at standard rates. The launch price became the standard price on August 10, 2026, and a planned increase was cancelled.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadGLM-5.3-FlashClaude Sonnet 5
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$0.16$2.40
Large one-off review, 150K input with no cache hits, 10K output$0.03$0.40
Output-heavy generation, 30K input, 80K output$0.04$0.86
A month of sessions, 110 sessions: 5 a day, 22 working days$17.60$264.00
Where the session’s cost goes
Cache writes$0.06$1.30
Cache reads$0.06$0.40
Uncached input$0.02$0.20
Output$0.03$0.50
caching saves on the session with GLM-5.3-Flash (60%)
$0.24
caching saves on the session with Claude Sonnet 5 (56%)
$3.10

What makes up a 15x gap

Every rate on GLM-5.3-Flash is a small fraction of its counterpart on Claude Sonnet 5. Input is $0.15 against $2 per million, 13.3x. Output is $0.50 against $10, 20x. A cache hit is $0.03 against $0.20, 6.7x, the closest of the four because Z.ai discounts a hit to 20% of input where Anthropic goes to 10%.

Cache writes separate them most in dollars. Z.ai lists no fee for writing its cache, so GLM-5.3-Flash's written tokens cost $0.15 per million. Anthropic charges 1.25x input for a 5-minute write and 2x for a 1-hour write, $2.50 and $4 on Sonnet 5. In the example session, which writes 400K tokens, that is $0.06 against $1.30, and the writes alone make up 54% of Sonnet 5's session cost.

The totals: $0.16 against $2.40 per session, 15x; $0.03 against $0.40 for the large one-off review; and $0.04 against $0.86 for the output-heavy generation, 21.5x. At 110 sessions a month, $17.60 against $264.00.

What the price gap does not tell you

A 15x price difference says nothing about how many turns a task takes or how much code has to be redone, and this page has no data on either. Anthropic says "Claude Sonnet 5 is built to be the most agentic Sonnet model yet" and puts its performance close to Claude Opus 4.8. Z.ai says GLM-5.3-Flash outperforms GLM-5.2 at one-tenth the price. Neither claim is a comparison with the other model.

Token counts differ as well. Sonnet 5's tokenizer counts about 1.0 to 1.35x as many tokens as Claude Sonnet 4.6 for the same text, and Z.ai tokenizes differently again, so one repository is not the same size on both. Sonnet 5 also runs adaptive thinking at high effort by default, which can raise output, the priciest line on its card.

The practical test is your own history. EveryToken prices Sonnet 5 at Anthropic's rates from Claude Code, Cursor, or OpenCode, and prices GLM-5.3-Flash when you use it through OpenRouter, so real sessions can be compared.

Claude Code, OpenCode, and open weights

Anthropic's Claude Code agent works with Claude models, and Sonnet 5 is also offered in Cursor, OpenRouter, OpenCode, and GitHub Copilot. GLM-5.3-Flash runs through OpenRouter and OpenCode, and OpenRouter's rankings show it as the most-used model for programming over the week before September 28, 2026, summed across nine languages.

On OpenRouter, a request for an open-weight model like GLM-5.3-Flash is served by one of several providers, and each may price it differently from Z.ai. The figures here use Z.ai's API prices and Anthropic's.

Z.ai publishes the weights under the MIT license, so GLM-5.3-Flash can run on your own hardware. Hosting costs depend on that hardware and how much of it you use, and they are not part of this comparison. Both models accept 1M tokens of context and write up to 128K per response.

Prompt caching

How each maker bills cached tokens

Z.ai

Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.

A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.

Source: Z.ai docs: Context caching

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

Your own numbers

See what GLM-5.3-Flash and Claude Sonnet 5 really cost you.

everyaitoken reads your OpenRouter, Claude Code, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

How much cheaper is GLM-5.3-Flash than Claude Sonnet 5?

13.3x on input and 20x on output per million tokens. The example cached session costs $0.16 against $2.40, 15x, and output-heavy generation costs $0.04 against $0.86.

Can I run GLM-5.3-Flash in Claude Code?

Claude Code is Anthropic's coding agent, built around Claude models and defaulting to them. With custom configuration it can reach other providers' compatible endpoints, which this post doesn't cover. GLM-5.3-Flash is offered in OpenCode, on OpenRouter, and on Z.ai's own API.

Do both models have the same context window?

Yes. Both accept 1M tokens and write up to 128K tokens per response.

Why is the gap bigger on output-heavy work?

Output is 20x apart per token, $0.50 against $10, while input is 13.3x apart. A job that writes 80K tokens from a 30K prompt leans on the wider ratio, so it comes to 21.5x.

  • Claude Fable 5.1 vs Claude Sonnet 5

    Claude Fable 5.1 lists at 5x the rates of Claude Sonnet 5, but cache hits cost $0.25 against $0.20. Why sessions cost 4.4x, and who Anthropic aims Fable 5.1 at.

  • Claude Opus 5.5 vs Claude Sonnet 5

    Claude Opus 5.5 costs twice as much per token as Claude Sonnet 5, yet cache hits cost $0.20 on both. What that means for agentic coding sessions.

  • Claude Sonnet 5 vs Claude Haiku 4.5

    Claude Haiku 4.5 costs half as much as Claude Sonnet 5, with a 200K context window and a retirement date ahead. When the cheaper model fits a coding workflow.

  • Claude Sonnet 5 vs Claude Sonnet 4.6

    Claude Sonnet 5 costs 33% less per token than Claude Sonnet 4.6, but its tokenizer counts up to 1.35x as many tokens. What the upgrade saves in practice.

  • GLM-5.3 vs Claude Sonnet 5

    GLM-5.3 undercuts Claude Sonnet 5 by 40% on an agentic coding session, $1.44 against $2.40. Most of it comes from Anthropic's cache-write premium.

  • GLM-5.3 vs GLM-5.3-Flash

    GLM-5.3-Flash costs about a ninth of GLM-5.3 at Z.ai's rates: $0.16 against $1.44 for an agentic coding session. Same limits, same caching, different jobs.