Skip to content

Model comparison

GLM-5.3 vs Claude Opus 5.5: a 3.1x gap on coding sessions

GLM-5.3 costs about a third of Claude Opus 5.5 on an agentic coding session, $1.44 against $4.40, yet its cache hits cost more. Where the gap comes from.

· Prices as of September 28, 2026

  • GLM-5.3

    Z.ai · Released August 14, 2026

    Z.ai's flagship for complex software engineering and long-horizon agent work, with open weights.

    GLM-5.3 facts and comparisons
  • Claude Opus 5.5

    Anthropic · Released September 22, 2026

    Anthropic's recommended starting model for most work, built for long-running agentic coding. Anthropic says it matches Claude Fable 5.1 on most work at a much lower price.

    Claude Opus 5.5 facts and comparisons

The short answer

GLM-5.3 is the cheaper model by a wide margin: the example agentic coding session costs $1.44 on it against $4.40 on Claude Opus 5.5, a 3.1x gap, and output-heavy work is 4.4x apart. Pick Opus 5.5 if you work in Claude Code, where it is the default, or want the model Anthropic recommends for long, sprawling jobs; pick GLM-5.3 if cost leads the decision and you want open weights you can reach through OpenRouter or OpenCode.

Choose GLM-5.3 if

  • Cost leads the decision: GLM-5.3 lists $1.40 input and $4.40 output per million tokens, against $4 and $20 on Opus 5.5.
  • You want open weights, which Z.ai publishes under its own GLM-5.3 license, so running the model yourself is an option.
  • You already work in OpenCode or route requests through OpenRouter, where GLM-5.3 is offered.
  • You would rather pay a subscription than per token: Z.ai also sells GLM-5.3 with its GLM Coding Plan.

Choose Claude Opus 5.5 if

  • You live in Claude Code, which starts Pro, Max, Team, and Enterprise users, and API-key users, on Opus 5.5.
  • You run long, sprawling jobs such as codebase-wide migrations and audits, which Anthropic names as a strength of Opus 5.5.
  • Your work mostly rereads a cached prefix, where an Opus 5.5 cache hit costs $0.20 per million against GLM-5.3's $0.26, though its writes and output still cost more.
  • You pick models inside Cursor or GitHub Copilot, which offer Opus 5.5 and do not offer GLM-5.3.

Side by side

Specs and prices

FactGLM-5.3Claude Opus 5.5
MakerZ.aiAnthropic
API model idglm-5.3claude-opus-5-5
ReleasedAugust 14, 2026September 22, 2026
StatusCurrentCurrent
Context window1M tokens1M tokens
Max output128K tokens128K tokens
Open weightsYesNo
Input, per 1M tokens$1.40$4
Cache hit, per 1M$0.26$0.20
Cache write, per 1M$1.40 (same as input)$5 (5-minute), $8 (1-hour)
Output, per 1M tokens$4.40$20
Runs inOpenCode and OpenRouterClaude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker (GLM-5.3: September 28, 2026; Claude Opus 5.5: September 26, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Claude Opus 5.5: The full 1M context window is billed at standard rates. Fast mode, a research preview on the Claude API, costs $8 input and $40 output per million tokens.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadGLM-5.3Claude Opus 5.5
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$1.44$4.40
Large one-off review, 150K input with no cache hits, 10K output$0.25$0.80
Output-heavy generation, 30K input, 80K output$0.39$1.72
A month of sessions, 110 sessions: 5 a day, 22 working days$158.40$484.00
Where the session’s cost goes
Cache writes$0.56$2.60
Cache reads$0.52$0.40
Uncached input$0.14$0.40
Output$0.22$1.00
caching saves on the session with GLM-5.3 (61%)
$2.28
caching saves on the session with Claude Opus 5.5 (60%)
$6.60

How much cheaper is GLM-5.3 than Claude Opus 5.5?

Per token, GLM-5.3 is cheaper on almost every line. It charges $1.40 per million input tokens and $4.40 per million output tokens. Claude Opus 5.5 charges $4 and $20, so input costs 2.9x as much on the Anthropic model and output 4.5x as much. How that plays out depends on how much a task reads versus writes.

Across the three example workloads the gap runs from 3.1x to 4.4x. The large one-off review, 150K tokens of uncached input and 10K of output, costs $0.25 on GLM-5.3 and $0.80 on Opus 5.5. The output-heavy generation, 30K in and 80K out, costs $0.39 against $1.72, the widest spread, because output is where the list prices differ most. The agentic coding session shows the narrowest gap, $1.44 against $4.40.

At 110 sessions a month that becomes $158.40 against $484.00, a difference of $325.60. These are API-equivalent estimates at each maker's published rates. The GLM-5.3 figures use Z.ai's own API price, and an OpenRouter provider serving the open weights can charge a different rate.

Why the cheaper model has the dearer cache hit

One rate runs the other way. Z.ai charges $0.26 per million cached tokens on GLM-5.3, which is 18.6% of its input price. Anthropic bills a hit on Opus 5.5 at 0.05x input, or $0.20, so Opus 5.5 reads from its cache 23% more cheaply. In the session, 2M cached tokens cost $0.52 on GLM-5.3 and $0.40 on Opus 5.5.

Cache writes pull much harder in the opposite direction. Z.ai lists no fee for writing the cache, so the session's 400K written tokens are billed as ordinary input, $0.56 in all. Anthropic charges 1.25x input for a 5-minute write and 2x for a 1-hour write, $5 and $8 per million on Opus 5.5, and the same 400K tokens cost $2.60. That single line is 59% of the Opus 5.5 session and accounts for $2.04 of the $2.96 gap.

So the mix of reads and writes sets the gap. A session that rereads a stable prefix many times leans on the one rate where Opus 5.5 is cheaper, while one that keeps writing fresh context leans on Anthropic's write premium. Measured against sending the same tokens uncached, caching saves 61% on GLM-5.3 and 60% on Opus 5.5. The cache walkthrough on our homepage goes through the same session's token mix line by line.

Open weights, Claude Code, and where each model runs

GLM-5.3 ships with open weights under Z.ai's own GLM-5.3 license, so a team can run it on its own hardware instead of paying per token. What that costs depends on the hardware and the load, and this page does not estimate it. Opus 5.5 has no open weights and is used through hosted services.

Opus 5.5 is the default model in Claude Code, and Cursor, OpenCode, OpenRouter, and GitHub Copilot offer it too. GLM-5.3 is offered through OpenRouter and in OpenCode. Z.ai says its API serves GLM-5.3 through endpoints in the OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages formats, which helps with clients built for those APIs. Claude Code and Codex, though, are built around their makers' own models, and reaching other providers from them takes custom configuration this post doesn't cover.

If you use both, EveryToken prices Opus 5.5 at Anthropic's rates from your Claude Code, Cursor, and OpenCode history, and prices GLM-5.3 from OpenRouter's catalog when you use it through OpenRouter. It does not price GLM-5.3 called directly on Z.ai's own API.

What Z.ai and Anthropic say each model is for

Z.ai describes GLM-5.3 as its "latest flagship model, delivering comprehensive advancements in complex software engineering and agent capabilities." It positions the model for complex software engineering and long-horizon agent work. Reasoning is always on, at low, high, or max.

Anthropic calls Opus 5.5 its recommended starting model for most work and says "it performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5." Anthropic lists agentic coding, computer use, and knowledge work as its strengths and says its output is more than 30% faster than Claude Opus 5. Adaptive thinking is always on, at medium effort by default.

Neither maker's pitch tells you which model suits your code. Tokenizers differ, so the same file counts as a different number of tokens on each, and Anthropic notes that its newer tokenizer counts about 30% more tokens than earlier Claude models for the same text. Reasoning level changes how much each model writes, too. Run both on a representative task and compare the cost of a finished change, not the cost of a token.

Prompt caching

How each maker bills cached tokens

Z.ai

Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.

A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.

Source: Z.ai docs: Context caching

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

Your own numbers

See what GLM-5.3 and Claude Opus 5.5 really cost you.

everyaitoken reads your OpenRouter, Claude Code, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is GLM-5.3 cheaper than Claude Opus 5.5?

Yes, on list prices and on every example workload. The agentic session costs $1.44 against $4.40, the uncached review $0.25 against $0.80, and the output-heavy generation $0.39 against $1.72. Opus 5.5 is lower on one rate, the cache hit, at $0.20 against $0.26 per million.

Where can I use GLM-5.3 for coding?

GLM-5.3 is offered through OpenRouter and in OpenCode, and Z.ai's own API accepts requests in the OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages formats. Claude Code is built around Anthropic's models and defaults to Opus 5.5; with custom configuration it can reach other providers' compatible endpoints, which this post doesn't cover.

Does Z.ai charge for GLM-5.3 cache writes?

No separate fee. Z.ai caches repeated context automatically, and written tokens cost the ordinary $1.40 input rate. Z.ai adds that storing cached input costs nothing for a limited time. Opus 5.5 charges $5 per million for 5-minute writes and $8 for 1-hour writes.

Is self-hosting GLM-5.3 cheaper than using the API?

That depends on your hardware and how heavily you would use GLM-5.3, and this page does not estimate it. The open weights make self-hosting possible under Z.ai's GLM-5.3 license. The GLM-5.3 prices here are Z.ai's own API rates, and OpenRouter providers serving the same weights can charge differently.

  • Claude Fable 5.1 vs Claude Opus 5.5

    Claude Fable 5.1 lists at 2.5x the rates of Claude Opus 5.5. What that premium buys, why cheap cache hits barely narrow it, and when Anthropic suggests it.

  • Claude Opus 5.5 vs Claude Opus 4.8

    Claude Opus 5.5 lists 20% below Claude Opus 4.8 and bills cache hits at $0.20, not $0.50. What staying costs, and when Anthropic still recommends Opus 4.8.

  • Claude Opus 5.5 vs Claude Haiku 4.5

    Claude Opus 5.5 lists at 4x the rates of Claude Haiku 4.5, yet a cached coding session costs 3.7x. Context and output limits, and Haiku 4.5's retirement date.

  • Claude Opus 5.5 vs Claude Opus 5

    Claude Opus 5.5 is 20% cheaper per token than Claude Opus 5, and cache hits cost 60% less. What that means for Claude Code sessions, and what stays the same.

  • Claude Opus 5.5 vs Claude Sonnet 5

    Claude Opus 5.5 costs twice as much per token as Claude Sonnet 5, yet cache hits cost $0.20 on both. What that means for agentic coding sessions.

  • GLM-5.3 vs Claude Sonnet 5

    GLM-5.3 undercuts Claude Sonnet 5 by 40% on an agentic coding session, $1.44 against $2.40. Most of it comes from Anthropic's cache-write premium.