Skip to content

DeepSeek

DeepSeek-V4-Pro: price, context window, and caching

DeepSeek's agent-focused large model, released in general availability in August 2026 with support for OpenAI's Responses API and Codex.

Released August 13, 2026 · Prices as of September 28, 2026

In DeepSeek’s words

“The GA version of DeepSeek V4 Pro greatly enhances agent capabilities, with particularly significant performance improvements in production environments.”

DeepSeek API: Change log

What DeepSeek says it’s good at

  • Native support for OpenAI's Responses API, adapted for Codex with one-click setup Source
  • Reasoning effort levels low, high, and max Source

Facts

Specs and prices

FactDeepSeek-V4-Pro
MakerDeepSeek
API model iddeepseek-v4-pro
ReleasedAugust 13, 2026
StatusCurrent
Context window1M tokens
Max output384K tokens
Open weightsYes
Input, per 1M tokens$1.32
Cache hit, per 1M$0.044
Cache write, per 1M$1.32 (same as input)
Output, per 1M tokens$3.96
Runs inOpenCode and OpenRouter

Standard API rates in US dollars, as published by the maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. DeepSeek-V4-Pro: Prices are DeepSeek's peak rates. Off-peak hours cost 50% less: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays.

Good to know

  • Open weights under the MIT license. The current snapshot is DeepSeek-V4-Pro-0813.
  • DeepSeek says service continues with billing unchanged until a V4.1 Pro model arrives.

Cost

What typical work costs

Example token counts at DeepSeek-V4-Pro’s published rates. On the agentic session, caching saves $2.55 against billing every token as ordinary input.

Example workload costs for DeepSeek-V4-Pro
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$0.95
Large one-off review, 150K input with no cache hits, 10K output$0.24
Output-heavy generation, 30K input, 80K output$0.36
A month of sessions, 110 sessions: 5 a day, 22 working days$104.06

Prompt caching

How DeepSeek bills cached tokens

DeepSeek

DeepSeek's disk cache is on by default for every account, with no code changes. DeepSeek lists no fee for writing the cache.

A cache hit costs $0.006 per million tokens on DeepSeek-V4.1-Flash and $0.044 on DeepSeek-V4-Pro at peak rates, and off-peak hours cost 50% less.

Source: DeepSeek API: Context caching

DeepSeek-V4-Pro compared

  • DeepSeek-V4-Pro vs Claude Opus 5.5

    Claude Opus 5.5 costs $4.40 for a cached coding session that costs $0.95 on DeepSeek-V4-Pro. Cache writes and output drive the gap; a V4.1 Pro is planned.

  • DeepSeek-V4-Pro vs GLM-5.3

    DeepSeek-V4-Pro and GLM-5.3 list input and output within 11% of each other, but a DeepSeek cache hit costs $0.044 against $0.26. What that does to a session.

  • DeepSeek-V4-Pro vs GPT-6 Sol

    DeepSeek-V4-Pro costs $0.95 for a cached coding session against $2.10 on GPT-6 Sol. How the gap forms, and what DeepSeek's Codex support does and doesn't mean.

  • DeepSeek-V4.1-Flash vs DeepSeek-V4-Pro

    DeepSeek says DeepSeek-V4.1-Flash outperforms its own DeepSeek-V4-Pro, which costs 4.3x as much per coding session. What Pro still offers, and what comes next.

  • DeepSeek-V4-Pro vs Claude Sonnet 5

    DeepSeek-V4-Pro costs $0.95 for a cached coding session that costs $2.40 on Claude Sonnet 5, and cache writes explain most of it. Rates, limits, and tools.

  • Kimi K3 vs DeepSeek-V4-Pro

    An agentic coding session costs $2.85 on Kimi K3 and $0.95 on DeepSeek-V4-Pro at peak rates, and DeepSeek's off-peak hours cost 50% less again.

Your own numbers

See what DeepSeek-V4-Pro really costs you.

everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math