Skip to content

Model comparison

DeepSeek-V4.1-Flash vs DeepSeek-V4-Pro: is Pro worth 4.3x?

DeepSeek says DeepSeek-V4.1-Flash outperforms its own DeepSeek-V4-Pro, which costs 4.3x as much per coding session. What Pro still offers, and what comes next.

· Prices as of September 28, 2026

  • DeepSeek-V4.1-Flash

    DeepSeek · Released September 10, 2026

    DeepSeek's small, low-cost, natively multimodal model on its new architecture, which DeepSeek says outperforms its own DeepSeek-V4-Pro.

    DeepSeek-V4.1-Flash facts and comparisons
  • DeepSeek-V4-Pro

    DeepSeek · Released August 13, 2026

    DeepSeek's agent-focused large model, released in general availability in August 2026 with support for OpenAI's Responses API and Codex.

    DeepSeek-V4-Pro facts and comparisons

The short answer

DeepSeek-V4.1-Flash costs $0.22 for the example agentic coding session against $0.95 on DeepSeek-V4-Pro, and DeepSeek itself says tests by several parties put V4.1 Flash ahead of V4-Pro on performance, cost, speed, and total runtime. DeepSeek-V4-Pro keeps a case for setups built on its Responses API support and Codex adaptation, or on its low, high, and max effort levels. DeepSeek says V4-Pro service continues, billed as today, until a V4.1 Pro model arrives.

Choose DeepSeek-V4.1-Flash if

  • You want the newer of the two, released on September 10, 2026, which DeepSeek says outperforms V4-Pro.
  • Price matters: $0.30 input and $1.20 output per million at peak, against $1.32 and $3.96.
  • You want image input, which DeepSeek describes as native visual understanding in V4.1 Flash.
  • Your code already calls deepseek-flash or the older deepseek-v4-flash ids, which now route to V4.1 Flash.

Choose DeepSeek-V4-Pro if

  • You rely on the Responses API support and one-click Codex setup that DeepSeek documents for V4-Pro.
  • You want reasoning effort you can set to low, high, or max, as DeepSeek lists for V4-Pro.
  • You want to stay on a fixed snapshot, DeepSeek-V4-Pro-0813, until DeepSeek ships V4.1 Pro.

Side by side

Specs and prices

FactDeepSeek-V4.1-FlashDeepSeek-V4-Pro
MakerDeepSeekDeepSeek
API model iddeepseek-flashdeepseek-v4-pro
ReleasedSeptember 10, 2026August 13, 2026
StatusCurrentCurrent
Context window1M tokens1M tokens
Max output384K tokens384K tokens
Open weightsYesYes
Input, per 1M tokens$0.30$1.32
Cache hit, per 1M$0.006$0.044
Cache write, per 1M$0.30 (same as input)$1.32 (same as input)
Output, per 1M tokens$1.20$3.96
Runs inOpenCode and OpenRouterOpenCode and OpenRouter

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. DeepSeek-V4.1-Flash: Prices are DeepSeek's peak rates. Off-peak hours cost 50% less: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. DeepSeek-V4-Pro: Prices are DeepSeek's peak rates. Off-peak hours cost 50% less: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadDeepSeek-V4.1-FlashDeepSeek-V4-Pro
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$0.22$0.95
Large one-off review, 150K input with no cache hits, 10K output$0.06$0.24
Output-heavy generation, 30K input, 80K output$0.11$0.36
A month of sessions, 110 sessions: 5 a day, 22 working days$24.42$104.06
Where the session’s cost goes
Cache writes$0.12$0.53
Cache reads$0.01$0.09
Uncached input$0.03$0.13
Output$0.06$0.20
caching saves on the session with DeepSeek-V4.1-Flash (73%)
$0.59
caching saves on the session with DeepSeek-V4-Pro (73%)
$2.55

DeepSeek's own answer points to the newer Flash model

DeepSeek's guidance favors the cheaper model here. Announcing DeepSeek-V4.1-Flash on September 10, 2026, DeepSeek said tests by several parties put it ahead of DeepSeek-V4-Pro on performance, cost, speed, and total runtime. It introduced V4.1 Flash as "the smallest model in our new architecture family, with native visual understanding."

V4-Pro reached general availability on August 13, 2026. DeepSeek said at the time: "The GA version of DeepSeek V4 Pro greatly enhances agent capabilities, with particularly significant performance improvements in production environments." The current snapshot is DeepSeek-V4-Pro-0813, and DeepSeek says it keeps serving V4-Pro at unchanged prices until a V4.1 Pro model arrives.

None of these claims are tested here. This page compares prices and published specs, and apart from price the specs match: both accept 1M tokens of context, both write up to 384K tokens, both publish open weights under the MIT license, and both follow DeepSeek's cache rules and off-peak discount.

How much more V4-Pro costs, line by line

V4-Pro lists $1.32 per million input tokens and $3.96 per million output tokens at peak, against $0.30 and $1.20 for V4.1 Flash: 4.4x on input and 3.3x on output. Cache hits cost $0.044 against $0.006, 7.3x apart, and DeepSeek charges nothing extra to write the cache on either model, so written tokens cost each one's input rate.

On the example session that comes to $0.95 against $0.22, 4.3x, or $104.06 against $24.42 over 110 sessions a month. The large one-off review is 4x, $0.24 against $0.06, and the output-heavy generation 3.3x, $0.36 against $0.11. Caching saves 73% of the uncached session cost on both, because both hit prices sit far below input: 3.3% of it on V4-Pro and 2% on V4.1 Flash.

The off-peak rule is shared. Peak runs from 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays, and every other hour costs 50% less on both models, so the ratio between them holds at any hour.

Reasons some teams stay on V4-Pro

DeepSeek built V4-Pro for agents that speak OpenAI's protocol. It says V4-Pro supports OpenAI's Responses API natively and is adapted for Codex, OpenAI's own coding agent, with one-click setup. It also lists three reasoning effort levels: low, high, and max. Those features are documented for V4-Pro, so confirm them for V4.1 Flash in DeepSeek's docs before moving a Codex setup.

Moving between them is mostly a change of model id. The id deepseek-flash serves V4.1 Flash, and the older deepseek-v4-flash ids route to it; V4-Pro keeps its own id, deepseek-v4-pro.

Both run through OpenRouter and OpenCode, and neither is in Cursor or GitHub Copilot. On OpenRouter a request goes to one of several providers, and their prices can differ from DeepSeek's own API prices used in these tables. By tokens processed, V4.1 Flash was the most-used model on OpenRouter in the week before September 28, 2026.

EveryToken prices either model when you run it through OpenRouter, from OpenRouter's catalog, so you can see how your own sessions split between them. It does not price calls made to DeepSeek's own API.

Prompt caching

How DeepSeek bills cached tokens

DeepSeek

DeepSeek's disk cache is on by default for every account, with no code changes. DeepSeek lists no fee for writing the cache.

A cache hit costs $0.006 per million tokens on DeepSeek-V4.1-Flash and $0.044 on DeepSeek-V4-Pro at peak rates, and off-peak hours cost 50% less.

Source: DeepSeek API: Context caching

Your own numbers

See what DeepSeek-V4.1-Flash and DeepSeek-V4-Pro really cost you.

everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is DeepSeek-V4.1-Flash better than DeepSeek-V4-Pro?

DeepSeek says so: it reports that tests by several parties put V4.1 Flash ahead of V4-Pro on performance, cost, speed, and total runtime. That is DeepSeek's claim, and this page does not test it.

How much cheaper is DeepSeek-V4.1-Flash?

At peak rates, input is $0.30 against $1.32 and output $1.20 against $3.96 per million. The example cached session costs $0.22 against $0.95, a 4.3x gap.

Will there be a DeepSeek V4.1 Pro?

DeepSeek says one is coming and that V4-Pro stays in service, at today's billing, until it arrives. No V4.1 Pro prices are in this comparison.

Do the two models share limits?

Yes. Both accept 1M tokens of context and write up to 384K tokens, and both publish open weights under the MIT license.

  • DeepSeek-V4-Pro vs Claude Opus 5.5

    Claude Opus 5.5 costs $4.40 for a cached coding session that costs $0.95 on DeepSeek-V4-Pro. Cache writes and output drive the gap; a V4.1 Pro is planned.

  • DeepSeek-V4-Pro vs GLM-5.3

    DeepSeek-V4-Pro and GLM-5.3 list input and output within 11% of each other, but a DeepSeek cache hit costs $0.044 against $0.26. What that does to a session.

  • DeepSeek-V4-Pro vs GPT-6 Sol

    DeepSeek-V4-Pro costs $0.95 for a cached coding session against $2.10 on GPT-6 Sol. How the gap forms, and what DeepSeek's Codex support does and doesn't mean.

  • DeepSeek-V4.1-Flash vs GPT-5.6 Luna

    DeepSeek-V4.1-Flash and GPT-5.6 Luna both price a cached coding session at $0.22, from opposite ends: Luna's cheaper input, DeepSeek's cheaper cache hits.

  • DeepSeek-V4.1-Flash vs Gemini 3.8 Flash

    DeepSeek-V4.1-Flash costs $0.22 for a cached coding session that costs $0.71 on Gemini 3.8 Flash, and Gemini's introductory rates end December 31, 2026.

  • DeepSeek-V4.1-Flash vs Claude Haiku 4.5

    DeepSeek-V4.1-Flash holds 5x the context of Claude Haiku 4.5 and costs $0.22 against $1.20 for a cached coding session. Haiku 4.5's successor is announced.