Skip to content

Model comparison

DeepSeek-V4.1-Flash vs GLM-5.3-Flash: two open Flash models

GLM-5.3-Flash lists half DeepSeek-V4.1-Flash's input price, but GLM's cache hits cost 5x as much. A cached coding session: $0.16 against $0.22 at peak.

· Prices as of September 28, 2026

  • DeepSeek-V4.1-Flash

    DeepSeek · Released September 10, 2026

    DeepSeek's small, low-cost, natively multimodal model on its new architecture, which DeepSeek says outperforms its own DeepSeek-V4-Pro.

    DeepSeek-V4.1-Flash facts and comparisons
  • GLM-5.3-Flash

    Z.ai · Released August 26, 2026

    Z.ai's low-cost, natively multimodal model, which Z.ai says outperforms GLM-5.2 at a tenth of the price.

    GLM-5.3-Flash facts and comparisons

The short answer

GLM-5.3-Flash has the lower list prices, and the example agentic coding session costs $0.16 on it against $0.22 on DeepSeek-V4.1-Flash, a gap held to 1.4x because GLM's cache hits cost 5x as much as DeepSeek's. Off-peak, DeepSeek-V4.1-Flash charges 50% less, which is enough to reverse the session result. Both are MIT-licensed open-weight models that run through OpenRouter and OpenCode, so the choice turns on when you work, how much you write, and how much you cache.

Choose DeepSeek-V4.1-Flash if

  • Your agent rereads a big cached context, and a DeepSeek hit costs $0.006 per million where GLM-5.3-Flash charges $0.03.
  • Most of your sessions run when DeepSeek is off-peak, and every rate is 50% lower.
  • You need up to 384K output tokens in one response, 3x GLM-5.3-Flash's 128K.
  • You are coming from DeepSeek-V4-Pro, which DeepSeek says V4.1 Flash beats on performance, cost, speed, and total runtime.

Choose GLM-5.3-Flash if

  • Your team works through DeepSeek's peak hours, when GLM-5.3-Flash's $0.15 input and $0.50 output undercut DeepSeek's $0.30 and $1.20.
  • Your work is output-heavy: the generation workload costs $0.04 here against $0.11.
  • You want a model that checks its own interface work, which Z.ai calls native multimodal visual coding.
  • Your requests are mostly fresh input with little cache reuse, where the full 2x input gap applies.

Side by side

Specs and prices

FactDeepSeek-V4.1-FlashGLM-5.3-Flash
MakerDeepSeekZ.ai
API model iddeepseek-flashglm-5.3-flash
ReleasedSeptember 10, 2026August 26, 2026
StatusCurrentCurrent
Context window1M tokens1M tokens
Max output384K tokens128K tokens
Open weightsYesYes
Input, per 1M tokens$0.30$0.15
Cache hit, per 1M$0.006$0.03
Cache write, per 1M$0.30 (same as input)$0.15 (same as input)
Output, per 1M tokens$1.20$0.50
Runs inOpenCode and OpenRouterOpenCode and OpenRouter

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. DeepSeek-V4.1-Flash: Prices are DeepSeek's peak rates. Off-peak hours cost 50% less: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadDeepSeek-V4.1-FlashGLM-5.3-Flash
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$0.22$0.16
Large one-off review, 150K input with no cache hits, 10K output$0.06$0.03
Output-heavy generation, 30K input, 80K output$0.11$0.04
A month of sessions, 110 sessions: 5 a day, 22 working days$24.42$17.60
Where the session’s cost goes
Cache writes$0.12$0.06
Cache reads$0.01$0.06
Uncached input$0.03$0.02
Output$0.06$0.03
caching saves on the session with DeepSeek-V4.1-Flash (73%)
$0.59
caching saves on the session with GLM-5.3-Flash (60%)
$0.24

Lower list prices against cheaper cache hits

GLM-5.3-Flash charges $0.15 per million input tokens and $0.50 per million output tokens on Z.ai's API. DeepSeek-V4.1-Flash charges $0.30 and $1.20 at DeepSeek's peak rates. On uncached work that means DeepSeek-V4.1-Flash costs 2x as much on the large one-off review, $0.06 against $0.03, and 2.8x as much on the output-heavy generation, $0.11 against $0.04.

Caching pulls the other way. Neither maker charges a fee to write the cache, so written tokens cost ordinary input on both. A hit, though, costs $0.006 per million on DeepSeek, 2% of input, and $0.03 on GLM-5.3-Flash, 20% of input. In the example session the 2M cached tokens cost $0.01 on DeepSeek and $0.06 on GLM-5.3-Flash, where reads are 35% of the session.

The session nets out at $0.16 on GLM-5.3-Flash and $0.22 on DeepSeek-V4.1-Flash, 1.4x, and 110 sessions a month come to $17.60 against $24.42. Caching trims 73% off DeepSeek's uncached session cost and 60% off GLM-5.3-Flash's.

DeepSeek's off-peak discount can flip the result

The DeepSeek figures are peak rates, in force from 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. In every other hour DeepSeek charges 50% less, and that includes cache hits.

At half price, the example session on DeepSeek-V4.1-Flash drops below GLM-5.3-Flash's $0.16. Uncached work does not flip: DeepSeek's discounted input matches GLM-5.3-Flash's $0.15 list price and its discounted output stays above $0.50, so large reviews end up about even and output-heavy jobs still favor GLM-5.3-Flash.

Those peak hours fall in the Beijing working day, 09:00 to 12:00 and 14:00 to 18:00 China time. Most of the working day in the Americas, and European afternoons, fall outside them.

Two MIT-licensed open-weight models compared

Both makers publish weights under the MIT license, so either can run on your own hardware; this page does not estimate hosting costs. Both accept 1M tokens of context and take visual input. DeepSeek introduces V4.1 Flash as "the smallest model in our new architecture family, with native visual understanding," and Z.ai describes GLM-5.3-Flash's visual coding as looking at interfaces and rendered results to test and improve its work.

Both makers point to memory-saving designs. DeepSeek says V4.1 Flash's KV cache needs a quarter of the memory of the previous generation, cutting cache-hit costs for agents. Z.ai says GLM-5.3-Flash's hybrid sparse and linear attention cuts attention compute and KV cache size.

They also lead different OpenRouter rankings. DeepSeek-V4.1-Flash was the most-used model on OpenRouter by tokens processed in the week before September 28, 2026, and GLM-5.3-Flash the most-used for programming over that same week, summed across nine languages. On OpenRouter a request for either goes to one of several providers, whose prices can differ from the makers' own API prices used here.

EveryToken prices both when you use them through OpenRouter, from OpenRouter's catalog. It does not price them when you call DeepSeek's or Z.ai's own APIs directly.

Prompt caching

How each maker bills cached tokens

DeepSeek

DeepSeek's disk cache is on by default for every account, with no code changes. DeepSeek lists no fee for writing the cache.

A cache hit costs $0.006 per million tokens on DeepSeek-V4.1-Flash and $0.044 on DeepSeek-V4-Pro at peak rates, and off-peak hours cost 50% less.

Source: DeepSeek API: Context caching

Z.ai

Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.

A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.

Source: Z.ai docs: Context caching

Your own numbers

See what DeepSeek-V4.1-Flash and GLM-5.3-Flash really cost you.

everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Which is cheaper, DeepSeek-V4.1-Flash or GLM-5.3-Flash?

At list prices, GLM-5.3-Flash: $0.15 input and $0.50 output per million against $0.30 and $1.20 at DeepSeek's peak. On the example cached session the gap is $0.16 against $0.22. Off-peak, DeepSeek charges 50% less and the session comes out cheaper on DeepSeek.

Do DeepSeek or Z.ai charge for cache writes?

Neither lists a fee for writing the cache, and both cache automatically. A hit costs $0.006 per million on DeepSeek-V4.1-Flash and $0.03 on GLM-5.3-Flash.

What are the output limits?

DeepSeek-V4.1-Flash lists 384K output tokens per response and GLM-5.3-Flash 128K. Both accept 1M tokens of context.

Are both models open source?

Both publish open weights under the MIT license. Self-hosting is possible for either, and what it costs depends on your hardware.

  • DeepSeek-V4.1-Flash vs DeepSeek-V4-Pro

    DeepSeek says DeepSeek-V4.1-Flash outperforms its own DeepSeek-V4-Pro, which costs 4.3x as much per coding session. What Pro still offers, and what comes next.

  • GLM-5.3 vs GLM-5.3-Flash

    GLM-5.3-Flash costs about a ninth of GLM-5.3 at Z.ai's rates: $0.16 against $1.44 for an agentic coding session. Same limits, same caching, different jobs.

  • GLM-5.3-Flash vs Claude Sonnet 5

    GLM-5.3-Flash costs $0.16 for a cached coding session that costs $2.40 on Claude Sonnet 5. What drives a 15x gap, and what price alone can't tell you.

  • DeepSeek-V4.1-Flash vs GPT-5.6 Luna

    DeepSeek-V4.1-Flash and GPT-5.6 Luna both price a cached coding session at $0.22, from opposite ends: Luna's cheaper input, DeepSeek's cheaper cache hits.

  • DeepSeek-V4.1-Flash vs Gemini 3.8 Flash

    DeepSeek-V4.1-Flash costs $0.22 for a cached coding session that costs $0.71 on Gemini 3.8 Flash, and Gemini's introductory rates end December 31, 2026.

  • DeepSeek-V4.1-Flash vs Claude Haiku 4.5

    DeepSeek-V4.1-Flash holds 5x the context of Claude Haiku 4.5 and costs $0.22 against $1.20 for a cached coding session. Haiku 4.5's successor is announced.