DeepSeek
DeepSeek-V4-Pro: price, context window, and caching
DeepSeek's agent-focused large model, released in general availability in August 2026 with support for OpenAI's Responses API and Codex.
Released August 13, 2026 · Prices as of September 28, 2026
In DeepSeek’s words
“The GA version of DeepSeek V4 Pro greatly enhances agent capabilities, with particularly significant performance improvements in production environments.”
Facts
Specs and prices
| Fact | DeepSeek-V4-Pro |
|---|---|
| Maker | DeepSeek |
| API model id | deepseek-v4-pro |
| Released | August 13, 2026 |
| Status | Current |
| Context window | 1M tokens |
| Max output | 384K tokens |
| Open weights | Yes |
| Input, per 1M tokens | $1.32 |
| Cache hit, per 1M | $0.044 |
| Cache write, per 1M | $1.32 (same as input) |
| Output, per 1M tokens | $3.96 |
| Runs in | OpenCode and OpenRouter |
Standard API rates in US dollars, as published by the maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. DeepSeek-V4-Pro: Prices are DeepSeek's peak rates. Off-peak hours cost 50% less: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays.
Good to know
- Open weights under the MIT license. The current snapshot is DeepSeek-V4-Pro-0813.
- DeepSeek says service continues with billing unchanged until a V4.1 Pro model arrives.
Cost
What typical work costs
Example token counts at DeepSeek-V4-Pro’s published rates. On the agentic session, caching saves $2.55 against billing every token as ordinary input.
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.95 |
|---|---|
| Large one-off review, 150K input with no cache hits, 10K output | $0.24 |
| Output-heavy generation, 30K input, 80K output | $0.36 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $104.06 |
Prompt caching
How DeepSeek bills cached tokens
DeepSeek
DeepSeek's disk cache is on by default for every account, with no code changes. DeepSeek lists no fee for writing the cache.
A cache hit costs $0.006 per million tokens on DeepSeek-V4.1-Flash and $0.044 on DeepSeek-V4-Pro at peak rates, and off-peak hours cost 50% less.
Source: DeepSeek API: Context caching
DeepSeek-V4-Pro compared
DeepSeek-V4-Pro vs Claude Opus 5.5
Claude Opus 5.5 costs $4.40 for a cached coding session that costs $0.95 on DeepSeek-V4-Pro. Cache writes and output drive the gap; a V4.1 Pro is planned.
DeepSeek-V4-Pro vs GLM-5.3
DeepSeek-V4-Pro and GLM-5.3 list input and output within 11% of each other, but a DeepSeek cache hit costs $0.044 against $0.26. What that does to a session.
DeepSeek-V4-Pro vs GPT-6 Sol
DeepSeek-V4-Pro costs $0.95 for a cached coding session against $2.10 on GPT-6 Sol. How the gap forms, and what DeepSeek's Codex support does and doesn't mean.
DeepSeek-V4.1-Flash vs DeepSeek-V4-Pro
DeepSeek says DeepSeek-V4.1-Flash outperforms its own DeepSeek-V4-Pro, which costs 4.3x as much per coding session. What Pro still offers, and what comes next.
DeepSeek-V4-Pro vs Claude Sonnet 5
DeepSeek-V4-Pro costs $0.95 for a cached coding session that costs $2.40 on Claude Sonnet 5, and cache writes explain most of it. Rates, limits, and tools.
Kimi K3 vs DeepSeek-V4-Pro
An agentic coding session costs $2.85 on Kimi K3 and $0.95 on DeepSeek-V4-Pro at peak rates, and DeepSeek's off-peak hours cost 50% less again.
Your own numbers
See what DeepSeek-V4-Pro really costs you.
everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.