Model comparison
DeepSeek-V4-Pro vs GLM-5.3: similar rates, a 5.9x cache gap
DeepSeek-V4-Pro and GLM-5.3 list input and output within 11% of each other, but a DeepSeek cache hit costs $0.044 against $0.26. What that does to a session.
· Prices as of September 28, 2026
DeepSeek-V4-Pro
DeepSeek · Released August 13, 2026
DeepSeek's agent-focused large model, released in general availability in August 2026 with support for OpenAI's Responses API and Codex.
DeepSeek-V4-Pro facts and comparisonsGLM-5.3
Z.ai · Released August 14, 2026
Z.ai's flagship for complex software engineering and long-horizon agent work, with open weights.
GLM-5.3 facts and comparisons
The short answer
DeepSeek-V4-Pro and GLM-5.3 list nearly the same input and output prices, but a DeepSeek cache hit costs $0.044 per million against $0.26, so the example agentic coding session costs $0.95 on DeepSeek-V4-Pro against $1.44 on GLM-5.3. For cached agentic work DeepSeek-V4-Pro is the cheaper pick, and off-peak it costs 50% less again; GLM-5.3 suits you if you want Z.ai's GLM Coding Plan or its Anthropic-format endpoint.
Choose DeepSeek-V4-Pro if
- Your agent rereads a long context: a DeepSeek-V4-Pro cache hit costs $0.044 per million, 83% less than GLM-5.3's $0.26.
- You can schedule work off-peak, when DeepSeek's prices are 50% lower.
- You need long outputs: DeepSeek-V4-Pro writes up to 384K tokens per request, 3x GLM-5.3's 128K.
- You want open weights under the MIT license, or an open model that DeepSeek says is adapted for Codex with one-click setup.
Choose GLM-5.3 if
- You would rather pay a subscription than per token: Z.ai sells GLM-5.3 with its GLM Coding Plan.
- Your client speaks the Anthropic Messages format, one of three API formats Z.ai serves GLM-5.3 on, alongside OpenAI Chat Completions and Responses.
Side by side
Specs and prices
| Fact | DeepSeek-V4-Pro | GLM-5.3 |
|---|---|---|
| Maker | DeepSeek | Z.ai |
| API model id | deepseek-v4-pro | glm-5.3 |
| Released | August 13, 2026 | August 14, 2026 |
| Status | Current | Current |
| Context window | 1M tokens | 1M tokens |
| Max output | 384K tokens | 128K tokens |
| Open weights | Yes | Yes |
| Input, per 1M tokens | $1.32 | $1.40 |
| Cache hit, per 1M | $0.044 | $0.26 |
| Cache write, per 1M | $1.32 (same as input) | $1.40 (same as input) |
| Output, per 1M tokens | $3.96 | $4.40 |
| Runs in | OpenCode and OpenRouter | OpenCode and OpenRouter |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. DeepSeek-V4-Pro: Prices are DeepSeek's peak rates. Off-peak hours cost 50% less: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | DeepSeek-V4-Pro | GLM-5.3 |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.95 | $1.44 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.24 | $0.25 |
| Output-heavy generation, 30K input, 80K output | $0.36 | $0.39 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $104.06 | $158.40 |
| Where the session’s cost goes | ||
| Cache writes | $0.53 | $0.56 |
| Cache reads | $0.09 | $0.52 |
| Uncached input | $0.13 | $0.14 |
| Output | $0.20 | $0.22 |
- caching saves on the session with DeepSeek-V4-Pro (73%)
- $2.55
- caching saves on the session with GLM-5.3 (61%)
- $2.28
Two open models released a day apart, priced almost alike
DeepSeek-V4-Pro reached general availability on August 13, 2026, and GLM-5.3 followed on August 14. Their list prices sit close together: $1.32 input and $3.96 output per million tokens for DeepSeek, $1.40 and $4.40 for GLM-5.3. Input differs by 6% and output by 10%.
On uncached work that closeness shows. The large one-off review costs $0.24 on DeepSeek-V4-Pro and $0.25 on GLM-5.3, and the output-heavy generation $0.36 against $0.39. For a one-off prompt, price barely separates them.
Both publish open weights, under different licenses: DeepSeek uses the MIT license, and Z.ai its own GLM-5.3 license. DeepSeek-V4-Pro and GLM-5.3 are both offered through OpenRouter and in OpenCode, and neither is in Cursor or GitHub Copilot.
The cache hit price is where they split
DeepSeek charges $0.044 per million for a cache hit on V4-Pro, 3.3% of its input price. Z.ai charges $0.26 on GLM-5.3, or 18.6% of input. A GLM-5.3 hit therefore costs 5.9x as much as a DeepSeek one, while every other rate is within 11%.
Agentic sessions are mostly cache reads, so the hit price reshapes the DeepSeek-V4-Pro and GLM-5.3 totals. The example session reads 2M tokens from the cache: $0.09 on DeepSeek-V4-Pro and $0.52 on GLM-5.3. That $0.43 is nearly all of the $0.49 session gap, and the session ends at $0.95 against $1.44, or $104.06 against $158.40 over 110 sessions a month.
Neither maker charges a premium to write the cache. DeepSeek's disk cache is on by default for every account with no code changes, and Z.ai caches repeated context automatically, with cached input storage free for a limited time. With writes billed as input on both, they make up 56% of the DeepSeek session and 39% of the GLM-5.3 one. Caching saves 73% on DeepSeek-V4-Pro and 61% on GLM-5.3 against sending the same tokens uncached, and the cache example on our homepage walks through it.
DeepSeek's off-peak hours and the coming V4.1 Pro
Every DeepSeek figure in this comparison uses peak rates, and off-peak hours cost 50% less. DeepSeek defines peak as 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays, so most of the week is off-peak. Overnight refactors and weekend runs can land at half price.
DeepSeek says V4-Pro service continues with billing unchanged until a V4.1 Pro model arrives, and the current snapshot is DeepSeek-V4-Pro-0813. DeepSeek also says its smaller DeepSeek-V4.1-Flash, built on a new architecture, outperforms V4-Pro. Anyone choosing V4-Pro as of September 2026 should expect a successor.
Z.ai positions GLM-5.3 as its flagship for complex software engineering and long-horizon agent work, and sells it through its GLM Coding Plan subscription as well as per token. GLM-5.3 always reasons, at low, high, or max, and DeepSeek-V4-Pro offers the same three effort levels.
Limits, API formats, and what each maker says
Both models take 1M tokens of context. DeepSeek-V4-Pro writes up to 384K tokens per request, 3x the 128K limit on GLM-5.3, which matters for agents that generate large files in a single turn.
DeepSeek says "the GA version of DeepSeek V4 Pro greatly enhances agent capabilities, with particularly significant performance improvements in production environments," and it has added native support for OpenAI's Responses API, adapted for Codex with one-click setup. Z.ai says GLM-5.3 delivers "comprehensive advancements in complex software engineering and agent capabilities" and serves it on OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages endpoints.
The prices in this comparison are each maker's own. Through OpenRouter, a DeepSeek-V4-Pro or GLM-5.3 request goes to one of several providers, whose prices can differ. EveryToken prices both models from OpenRouter's catalog when you use them through OpenRouter, and does not price calls made directly to DeepSeek's or Z.ai's APIs.
Prompt caching
How each maker bills cached tokens
DeepSeek
DeepSeek's disk cache is on by default for every account, with no code changes. DeepSeek lists no fee for writing the cache.
A cache hit costs $0.006 per million tokens on DeepSeek-V4.1-Flash and $0.044 on DeepSeek-V4-Pro at peak rates, and off-peak hours cost 50% less.
Source: DeepSeek API: Context caching
Z.ai
Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.
A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.
Source: Z.ai docs: Context caching
Your own numbers
See what DeepSeek-V4-Pro and GLM-5.3 really cost you.
everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is DeepSeek-V4-Pro cheaper than GLM-5.3?
Slightly on list prices and by a lot on cache hits. Uncached work costs about the same, $0.24 against $0.25 for the large review, but the cached agentic session costs $0.95 against $1.44 because a DeepSeek hit is $0.044 per million against $0.26.
When are DeepSeek's off-peak hours?
DeepSeek's peak window is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, Chinese public holidays excepted. Outside those hours DeepSeek's prices are 50% lower. The figures on this page use the peak rates.
Do both models have open weights?
Yes. DeepSeek-V4-Pro is under the MIT license, and GLM-5.3 under Z.ai's own GLM-5.3 license. Either DeepSeek-V4-Pro or GLM-5.3 can be self-hosted, and this page does not estimate what that costs.
Is DeepSeek replacing V4-Pro?
DeepSeek says service continues with billing unchanged until a V4.1 Pro model arrives. It has already released DeepSeek-V4.1-Flash, which it says outperforms V4-Pro.