Model comparison
DeepSeek-V4.1-Flash vs GLM-5.3-Flash: two open Flash models
GLM-5.3-Flash lists half DeepSeek-V4.1-Flash's input price, but GLM's cache hits cost 5x as much. A cached coding session: $0.16 against $0.22 at peak.
· Prices as of September 28, 2026
DeepSeek-V4.1-Flash
DeepSeek · Released September 10, 2026
DeepSeek's small, low-cost, natively multimodal model on its new architecture, which DeepSeek says outperforms its own DeepSeek-V4-Pro.
DeepSeek-V4.1-Flash facts and comparisonsGLM-5.3-Flash
Z.ai · Released August 26, 2026
Z.ai's low-cost, natively multimodal model, which Z.ai says outperforms GLM-5.2 at a tenth of the price.
GLM-5.3-Flash facts and comparisons
The short answer
GLM-5.3-Flash has the lower list prices, and the example agentic coding session costs $0.16 on it against $0.22 on DeepSeek-V4.1-Flash, a gap held to 1.4x because GLM's cache hits cost 5x as much as DeepSeek's. Off-peak, DeepSeek-V4.1-Flash charges 50% less, which is enough to reverse the session result. Both are MIT-licensed open-weight models that run through OpenRouter and OpenCode, so the choice turns on when you work, how much you write, and how much you cache.
Choose DeepSeek-V4.1-Flash if
- Your agent rereads a big cached context, and a DeepSeek hit costs $0.006 per million where GLM-5.3-Flash charges $0.03.
- Most of your sessions run when DeepSeek is off-peak, and every rate is 50% lower.
- You need up to 384K output tokens in one response, 3x GLM-5.3-Flash's 128K.
- You are coming from DeepSeek-V4-Pro, which DeepSeek says V4.1 Flash beats on performance, cost, speed, and total runtime.
Choose GLM-5.3-Flash if
- Your team works through DeepSeek's peak hours, when GLM-5.3-Flash's $0.15 input and $0.50 output undercut DeepSeek's $0.30 and $1.20.
- Your work is output-heavy: the generation workload costs $0.04 here against $0.11.
- You want a model that checks its own interface work, which Z.ai calls native multimodal visual coding.
- Your requests are mostly fresh input with little cache reuse, where the full 2x input gap applies.
Side by side
Specs and prices
| Fact | DeepSeek-V4.1-Flash | GLM-5.3-Flash |
|---|---|---|
| Maker | DeepSeek | Z.ai |
| API model id | deepseek-flash | glm-5.3-flash |
| Released | September 10, 2026 | August 26, 2026 |
| Status | Current | Current |
| Context window | 1M tokens | 1M tokens |
| Max output | 384K tokens | 128K tokens |
| Open weights | Yes | Yes |
| Input, per 1M tokens | $0.30 | $0.15 |
| Cache hit, per 1M | $0.006 | $0.03 |
| Cache write, per 1M | $0.30 (same as input) | $0.15 (same as input) |
| Output, per 1M tokens | $1.20 | $0.50 |
| Runs in | OpenCode and OpenRouter | OpenCode and OpenRouter |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. DeepSeek-V4.1-Flash: Prices are DeepSeek's peak rates. Off-peak hours cost 50% less: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | DeepSeek-V4.1-Flash | GLM-5.3-Flash |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.22 | $0.16 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.06 | $0.03 |
| Output-heavy generation, 30K input, 80K output | $0.11 | $0.04 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $24.42 | $17.60 |
| Where the session’s cost goes | ||
| Cache writes | $0.12 | $0.06 |
| Cache reads | $0.01 | $0.06 |
| Uncached input | $0.03 | $0.02 |
| Output | $0.06 | $0.03 |
- caching saves on the session with DeepSeek-V4.1-Flash (73%)
- $0.59
- caching saves on the session with GLM-5.3-Flash (60%)
- $0.24
Lower list prices against cheaper cache hits
GLM-5.3-Flash charges $0.15 per million input tokens and $0.50 per million output tokens on Z.ai's API. DeepSeek-V4.1-Flash charges $0.30 and $1.20 at DeepSeek's peak rates. On uncached work that means DeepSeek-V4.1-Flash costs 2x as much on the large one-off review, $0.06 against $0.03, and 2.8x as much on the output-heavy generation, $0.11 against $0.04.
Caching pulls the other way. Neither maker charges a fee to write the cache, so written tokens cost ordinary input on both. A hit, though, costs $0.006 per million on DeepSeek, 2% of input, and $0.03 on GLM-5.3-Flash, 20% of input. In the example session the 2M cached tokens cost $0.01 on DeepSeek and $0.06 on GLM-5.3-Flash, where reads are 35% of the session.
The session nets out at $0.16 on GLM-5.3-Flash and $0.22 on DeepSeek-V4.1-Flash, 1.4x, and 110 sessions a month come to $17.60 against $24.42. Caching trims 73% off DeepSeek's uncached session cost and 60% off GLM-5.3-Flash's.
DeepSeek's off-peak discount can flip the result
The DeepSeek figures are peak rates, in force from 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. In every other hour DeepSeek charges 50% less, and that includes cache hits.
At half price, the example session on DeepSeek-V4.1-Flash drops below GLM-5.3-Flash's $0.16. Uncached work does not flip: DeepSeek's discounted input matches GLM-5.3-Flash's $0.15 list price and its discounted output stays above $0.50, so large reviews end up about even and output-heavy jobs still favor GLM-5.3-Flash.
Those peak hours fall in the Beijing working day, 09:00 to 12:00 and 14:00 to 18:00 China time. Most of the working day in the Americas, and European afternoons, fall outside them.
Two MIT-licensed open-weight models compared
Both makers publish weights under the MIT license, so either can run on your own hardware; this page does not estimate hosting costs. Both accept 1M tokens of context and take visual input. DeepSeek introduces V4.1 Flash as "the smallest model in our new architecture family, with native visual understanding," and Z.ai describes GLM-5.3-Flash's visual coding as looking at interfaces and rendered results to test and improve its work.
Both makers point to memory-saving designs. DeepSeek says V4.1 Flash's KV cache needs a quarter of the memory of the previous generation, cutting cache-hit costs for agents. Z.ai says GLM-5.3-Flash's hybrid sparse and linear attention cuts attention compute and KV cache size.
They also lead different OpenRouter rankings. DeepSeek-V4.1-Flash was the most-used model on OpenRouter by tokens processed in the week before September 28, 2026, and GLM-5.3-Flash the most-used for programming over that same week, summed across nine languages. On OpenRouter a request for either goes to one of several providers, whose prices can differ from the makers' own API prices used here.
EveryToken prices both when you use them through OpenRouter, from OpenRouter's catalog. It does not price them when you call DeepSeek's or Z.ai's own APIs directly.
Prompt caching
How each maker bills cached tokens
DeepSeek
DeepSeek's disk cache is on by default for every account, with no code changes. DeepSeek lists no fee for writing the cache.
A cache hit costs $0.006 per million tokens on DeepSeek-V4.1-Flash and $0.044 on DeepSeek-V4-Pro at peak rates, and off-peak hours cost 50% less.
Source: DeepSeek API: Context caching
Z.ai
Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.
A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.
Source: Z.ai docs: Context caching
Your own numbers
See what DeepSeek-V4.1-Flash and GLM-5.3-Flash really cost you.
everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Which is cheaper, DeepSeek-V4.1-Flash or GLM-5.3-Flash?
At list prices, GLM-5.3-Flash: $0.15 input and $0.50 output per million against $0.30 and $1.20 at DeepSeek's peak. On the example cached session the gap is $0.16 against $0.22. Off-peak, DeepSeek charges 50% less and the session comes out cheaper on DeepSeek.
Do DeepSeek or Z.ai charge for cache writes?
Neither lists a fee for writing the cache, and both cache automatically. A hit costs $0.006 per million on DeepSeek-V4.1-Flash and $0.03 on GLM-5.3-Flash.
What are the output limits?
DeepSeek-V4.1-Flash lists 384K output tokens per response and GLM-5.3-Flash 128K. Both accept 1M tokens of context.
Are both models open source?
Both publish open weights under the MIT license. Self-hosting is possible for either, and what it costs depends on your hardware.