Skip to content

Model comparison

GLM-5.3-Flash vs Claude Haiku 4.5: the 10x output gap

GLM-5.3-Flash charges $0.50 per million output tokens against $5 on Claude Haiku 4.5, and $0.16 against $1.20 for a cached coding session. Where each fits.

· Prices as of September 28, 2026

  • GLM-5.3-Flash

    Z.ai · Released August 26, 2026

    Z.ai's low-cost, natively multimodal model, which Z.ai says outperforms GLM-5.2 at a tenth of the price.

    GLM-5.3-Flash facts and comparisons
  • Claude Haiku 4.5

    Anthropic · Released October 15, 2025

    Anthropic's fastest and cheapest current model, aimed at latency-sensitive work and at coding sub-agents.

    Claude Haiku 4.5 facts and comparisons

The short answer

GLM-5.3-Flash costs $0.16 for the example agentic coding session against $1.20 on Claude Haiku 4.5, and the gap is widest on output-heavy work, 10.8x, since Z.ai charges $0.50 per million output tokens against $5. Claude Haiku 4.5 fits Claude Code users who want Anthropic's low-cost model for sub-agents, though Anthropic has announced Claude Haiku 5.5 and lists Haiku 4.5's retirement as not sooner than October 15, 2026. GLM-5.3-Flash suits high-volume work through OpenRouter or OpenCode, with 1M tokens of context and MIT-licensed weights.

Choose GLM-5.3-Flash if

  • Your jobs write a lot: GLM-5.3-Flash charges $0.50 per million output tokens, a tenth of Haiku 4.5's $5.
  • Your prompts can pass 200K tokens, which GLM-5.3-Flash handles with its 1M window.
  • You want a model that inspects what it builds, which Z.ai describes as native multimodal visual coding.
  • You want MIT-licensed open weights that can run on your own hardware.

Choose Claude Haiku 4.5 if

  • You use Claude Code and want Anthropic's low-cost model for sub-agents in multi-agent refactors and migrations.
  • Cursor or GitHub Copilot is where you code, and both offer Haiku 4.5 while neither lists GLM-5.3-Flash.
  • You prefer manual extended thinking with a token budget you control.

Side by side

Specs and prices

FactGLM-5.3-FlashClaude Haiku 4.5
MakerZ.aiAnthropic
API model idglm-5.3-flashclaude-haiku-4-5
ReleasedAugust 26, 2026October 15, 2025
StatusCurrentCurrent
Context window1M tokens200K tokens
Max output128K tokens64K tokens
Open weightsYesNo
Input, per 1M tokens$0.15$1
Cache hit, per 1M$0.03$0.10
Cache write, per 1M$0.15 (same as input)$1.25 (5-minute), $2 (1-hour)
Output, per 1M tokens$0.50$5
Runs inOpenCode and OpenRouterClaude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker (GLM-5.3-Flash: September 28, 2026; Claude Haiku 4.5: September 26, 2026). Batch and priority tiers, taxes, and subscription plans are not included.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadGLM-5.3-FlashClaude Haiku 4.5
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$0.16$1.20
Large one-off review, 150K input with no cache hits, 10K output$0.03$0.20
Output-heavy generation, 30K input, 80K output$0.04$0.43
A month of sessions, 110 sessions: 5 a day, 22 working days$17.60$132.00
Where the session’s cost goes
Cache writes$0.06$0.65
Cache reads$0.06$0.20
Uncached input$0.02$0.10
Output$0.03$0.25
caching saves on the session with GLM-5.3-Flash (60%)
$0.24
caching saves on the session with Claude Haiku 4.5 (56%)
$1.55

Output-heavy work shows the widest gap

GLM-5.3-Flash lists $0.15 per million input tokens and $0.50 per million output tokens on Z.ai's API. Claude Haiku 4.5 lists $1 and $5. Input differs by 6.7x and output by 10x, so the more a job writes, the further apart the two land.

The output-heavy generation shows it: $0.04 on GLM-5.3-Flash against $0.43 on Haiku 4.5, 10.8x. The large one-off review, mostly input, sits at the input ratio, $0.03 against $0.20. Code generation, test writing, and long explanations sit toward the output end of that range; reviews and searches over a codebase sit toward the input end.

Z.ai's cache discount is shallower than Anthropic's

Z.ai caches repeated context automatically and lists no fee for writing the cache, so written tokens cost GLM-5.3-Flash's ordinary input rate. A hit costs $0.03 per million, 20% of input. Anthropic discounts a hit more steeply, to 10% of input or $0.10 on Haiku 4.5, but charges 1.25x input for a 5-minute write and 2x for a 1-hour write.

On the example session, those rules shape the two bills differently. Cache reads cost $0.06 on GLM-5.3-Flash, 35% of its session, against $0.20 on Haiku. Cache writes cost $0.06 against $0.65. The session totals $0.16 against $1.20, 7.5x, which sits between the review's 6.7x and the generation's 10.8x.

Caching saves 60% of the uncached session cost on GLM-5.3-Flash and 56% on Haiku 4.5, and Z.ai says cached input storage is free for a limited time. Over 110 sessions a month the totals are $17.60 and $132.00. The cache walkthrough on the homepage shows the same session line by line.

Context, retirement, and where each model runs

Haiku 4.5 accepts 200K tokens of context and writes up to 64K per response. GLM-5.3-Flash accepts 1M and writes up to 128K. Haiku 4.5 also uses Anthropic's older tokenizer, which counts fewer tokens for the same text than Claude models from 4.7 on, and Z.ai counts tokens its own way, so equal prompts will not show equal counts.

Anthropic lists Haiku 4.5's retirement as not sooner than October 15, 2026, and by September 28, 2026 it had announced Claude Haiku 5.5, due within weeks. It credits Haiku 4.5 with coding performance similar to Claude Sonnet 4 at one-third the cost and more than twice the speed. Z.ai, for its part, says GLM-5.3-Flash beats GLM-5.2 at a tenth of the price, and that its hybrid sparse and linear attention cuts attention compute and KV cache size.

Anthropic's own agent, Claude Code, runs Haiku 4.5 alongside the other Claude models, and Haiku 4.5 is also in Cursor, OpenRouter, OpenCode, and GitHub Copilot. GLM-5.3-Flash runs through OpenRouter and OpenCode, and it was the most-used model for programming on OpenRouter over the week before September 28, 2026, summed across nine languages. On OpenRouter, requests for open-weight models go to one of several providers, whose prices can differ from Z.ai's own, which these tables use.

EveryToken prices Haiku 4.5 at Anthropic's rates in Claude Code, Cursor, and OpenCode, and prices GLM-5.3-Flash when you use it through OpenRouter, from OpenRouter's catalog.

Prompt caching

How each maker bills cached tokens

Z.ai

Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.

A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.

Source: Z.ai docs: Context caching

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

Your own numbers

See what GLM-5.3-Flash and Claude Haiku 4.5 really cost you.

everyaitoken reads your OpenRouter, Claude Code, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

How much cheaper is GLM-5.3-Flash than Claude Haiku 4.5?

6.7x on input and 10x on output at Z.ai's list prices. The example cached session costs $0.16 against $1.20, and 110 sessions a month come to $17.60 against $132.00.

Does GLM-5.3-Flash charge for cache writes?

Z.ai lists no fee for writing the cache, and caching is automatic. A GLM-5.3-Flash cache hit costs $0.03 per million, 20% of input. Anthropic charges 1.25x or 2x input to write Haiku 4.5's cache, depending on the lifetime.

Is Claude Haiku 4.5 being retired?

Yes, eventually. Anthropic lists its retirement as not sooner than October 15, 2026, and as of late September 2026 it had announced Claude Haiku 5.5 as the next Haiku.

Can GLM-5.3-Flash look at screenshots of my app?

Z.ai describes native multimodal visual coding: the model looks at interfaces and rendered results to test and improve its work. Whether your coding tool passes images to it depends on the tool.

  • Claude Opus 5.5 vs Claude Haiku 4.5

    Claude Opus 5.5 lists at 4x the rates of Claude Haiku 4.5, yet a cached coding session costs 3.7x. Context and output limits, and Haiku 4.5's retirement date.

  • Claude Sonnet 5 vs Claude Haiku 4.5

    Claude Haiku 4.5 costs half as much as Claude Sonnet 5, with a 200K context window and a retirement date ahead. When the cheaper model fits a coding workflow.

  • GLM-5.3-Flash vs Claude Sonnet 5

    GLM-5.3-Flash costs $0.16 for a cached coding session that costs $2.40 on Claude Sonnet 5. What drives a 15x gap, and what price alone can't tell you.

  • GLM-5.3 vs GLM-5.3-Flash

    GLM-5.3-Flash costs about a ninth of GLM-5.3 at Z.ai's rates: $0.16 against $1.44 for an agentic coding session. Same limits, same caching, different jobs.

  • Claude Haiku 4.5 vs Gemini 3.5 Flash-Lite

    Claude Haiku 4.5 and Gemini 3.5 Flash-Lite both target sub-agent work. Haiku 4.5 costs 3.5x as much per cached session, but only 2x as much per output token.

  • Claude Haiku 4.5 vs Gemini 3.8 Flash

    Gemini 3.8 Flash undercuts Claude Haiku 4.5 on introductory rates, $0.71 against $1.20 per cached session. From January 1, 2027, its rates top Haiku's.