Model comparison
DeepSeek-V4.1-Flash vs DeepSeek-V4-Pro: is Pro worth 4.3x?
DeepSeek says DeepSeek-V4.1-Flash outperforms its own DeepSeek-V4-Pro, which costs 4.3x as much per coding session. What Pro still offers, and what comes next.
· Prices as of September 28, 2026
DeepSeek-V4.1-Flash
DeepSeek · Released September 10, 2026
DeepSeek's small, low-cost, natively multimodal model on its new architecture, which DeepSeek says outperforms its own DeepSeek-V4-Pro.
DeepSeek-V4.1-Flash facts and comparisonsDeepSeek-V4-Pro
DeepSeek · Released August 13, 2026
DeepSeek's agent-focused large model, released in general availability in August 2026 with support for OpenAI's Responses API and Codex.
DeepSeek-V4-Pro facts and comparisons
The short answer
DeepSeek-V4.1-Flash costs $0.22 for the example agentic coding session against $0.95 on DeepSeek-V4-Pro, and DeepSeek itself says tests by several parties put V4.1 Flash ahead of V4-Pro on performance, cost, speed, and total runtime. DeepSeek-V4-Pro keeps a case for setups built on its Responses API support and Codex adaptation, or on its low, high, and max effort levels. DeepSeek says V4-Pro service continues, billed as today, until a V4.1 Pro model arrives.
Choose DeepSeek-V4.1-Flash if
- You want the newer of the two, released on September 10, 2026, which DeepSeek says outperforms V4-Pro.
- Price matters: $0.30 input and $1.20 output per million at peak, against $1.32 and $3.96.
- You want image input, which DeepSeek describes as native visual understanding in V4.1 Flash.
- Your code already calls deepseek-flash or the older deepseek-v4-flash ids, which now route to V4.1 Flash.
Choose DeepSeek-V4-Pro if
- You rely on the Responses API support and one-click Codex setup that DeepSeek documents for V4-Pro.
- You want reasoning effort you can set to low, high, or max, as DeepSeek lists for V4-Pro.
- You want to stay on a fixed snapshot, DeepSeek-V4-Pro-0813, until DeepSeek ships V4.1 Pro.
Side by side
Specs and prices
| Fact | DeepSeek-V4.1-Flash | DeepSeek-V4-Pro |
|---|---|---|
| Maker | DeepSeek | DeepSeek |
| API model id | deepseek-flash | deepseek-v4-pro |
| Released | September 10, 2026 | August 13, 2026 |
| Status | Current | Current |
| Context window | 1M tokens | 1M tokens |
| Max output | 384K tokens | 384K tokens |
| Open weights | Yes | Yes |
| Input, per 1M tokens | $0.30 | $1.32 |
| Cache hit, per 1M | $0.006 | $0.044 |
| Cache write, per 1M | $0.30 (same as input) | $1.32 (same as input) |
| Output, per 1M tokens | $1.20 | $3.96 |
| Runs in | OpenCode and OpenRouter | OpenCode and OpenRouter |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. DeepSeek-V4.1-Flash: Prices are DeepSeek's peak rates. Off-peak hours cost 50% less: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. DeepSeek-V4-Pro: Prices are DeepSeek's peak rates. Off-peak hours cost 50% less: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | DeepSeek-V4.1-Flash | DeepSeek-V4-Pro |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.22 | $0.95 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.06 | $0.24 |
| Output-heavy generation, 30K input, 80K output | $0.11 | $0.36 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $24.42 | $104.06 |
| Where the session’s cost goes | ||
| Cache writes | $0.12 | $0.53 |
| Cache reads | $0.01 | $0.09 |
| Uncached input | $0.03 | $0.13 |
| Output | $0.06 | $0.20 |
- caching saves on the session with DeepSeek-V4.1-Flash (73%)
- $0.59
- caching saves on the session with DeepSeek-V4-Pro (73%)
- $2.55
DeepSeek's own answer points to the newer Flash model
DeepSeek's guidance favors the cheaper model here. Announcing DeepSeek-V4.1-Flash on September 10, 2026, DeepSeek said tests by several parties put it ahead of DeepSeek-V4-Pro on performance, cost, speed, and total runtime. It introduced V4.1 Flash as "the smallest model in our new architecture family, with native visual understanding."
V4-Pro reached general availability on August 13, 2026. DeepSeek said at the time: "The GA version of DeepSeek V4 Pro greatly enhances agent capabilities, with particularly significant performance improvements in production environments." The current snapshot is DeepSeek-V4-Pro-0813, and DeepSeek says it keeps serving V4-Pro at unchanged prices until a V4.1 Pro model arrives.
None of these claims are tested here. This page compares prices and published specs, and apart from price the specs match: both accept 1M tokens of context, both write up to 384K tokens, both publish open weights under the MIT license, and both follow DeepSeek's cache rules and off-peak discount.
How much more V4-Pro costs, line by line
V4-Pro lists $1.32 per million input tokens and $3.96 per million output tokens at peak, against $0.30 and $1.20 for V4.1 Flash: 4.4x on input and 3.3x on output. Cache hits cost $0.044 against $0.006, 7.3x apart, and DeepSeek charges nothing extra to write the cache on either model, so written tokens cost each one's input rate.
On the example session that comes to $0.95 against $0.22, 4.3x, or $104.06 against $24.42 over 110 sessions a month. The large one-off review is 4x, $0.24 against $0.06, and the output-heavy generation 3.3x, $0.36 against $0.11. Caching saves 73% of the uncached session cost on both, because both hit prices sit far below input: 3.3% of it on V4-Pro and 2% on V4.1 Flash.
The off-peak rule is shared. Peak runs from 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays, and every other hour costs 50% less on both models, so the ratio between them holds at any hour.
Reasons some teams stay on V4-Pro
DeepSeek built V4-Pro for agents that speak OpenAI's protocol. It says V4-Pro supports OpenAI's Responses API natively and is adapted for Codex, OpenAI's own coding agent, with one-click setup. It also lists three reasoning effort levels: low, high, and max. Those features are documented for V4-Pro, so confirm them for V4.1 Flash in DeepSeek's docs before moving a Codex setup.
Moving between them is mostly a change of model id. The id deepseek-flash serves V4.1 Flash, and the older deepseek-v4-flash ids route to it; V4-Pro keeps its own id, deepseek-v4-pro.
Both run through OpenRouter and OpenCode, and neither is in Cursor or GitHub Copilot. On OpenRouter a request goes to one of several providers, and their prices can differ from DeepSeek's own API prices used in these tables. By tokens processed, V4.1 Flash was the most-used model on OpenRouter in the week before September 28, 2026.
EveryToken prices either model when you run it through OpenRouter, from OpenRouter's catalog, so you can see how your own sessions split between them. It does not price calls made to DeepSeek's own API.
Prompt caching
How DeepSeek bills cached tokens
DeepSeek
DeepSeek's disk cache is on by default for every account, with no code changes. DeepSeek lists no fee for writing the cache.
A cache hit costs $0.006 per million tokens on DeepSeek-V4.1-Flash and $0.044 on DeepSeek-V4-Pro at peak rates, and off-peak hours cost 50% less.
Source: DeepSeek API: Context caching
Your own numbers
See what DeepSeek-V4.1-Flash and DeepSeek-V4-Pro really cost you.
everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is DeepSeek-V4.1-Flash better than DeepSeek-V4-Pro?
DeepSeek says so: it reports that tests by several parties put V4.1 Flash ahead of V4-Pro on performance, cost, speed, and total runtime. That is DeepSeek's claim, and this page does not test it.
How much cheaper is DeepSeek-V4.1-Flash?
At peak rates, input is $0.30 against $1.32 and output $1.20 against $3.96 per million. The example cached session costs $0.22 against $0.95, a 4.3x gap.
Will there be a DeepSeek V4.1 Pro?
DeepSeek says one is coming and that V4-Pro stays in service, at today's billing, until it arrives. No V4.1 Pro prices are in this comparison.
Do the two models share limits?
Yes. Both accept 1M tokens of context and write up to 384K tokens, and both publish open weights under the MIT license.