Skip to content

DeepSeek

DeepSeek-V4.1-Flash: price, context window, and caching

DeepSeek's small, low-cost, natively multimodal model on its new architecture, which DeepSeek says outperforms its own DeepSeek-V4-Pro.

Released September 10, 2026 · Prices as of September 28, 2026

In DeepSeek’s words

“Introducing the smallest model in our new architecture family, with native visual understanding.”

DeepSeek: V4.1 Flash release

What DeepSeek says it’s good at

  • A KV cache that needs a quarter of the memory of the previous generation, cutting cache-hit costs for agents Source
  • Ahead of DeepSeek-V4-Pro on performance, cost, speed, and total runtime in tests by several parties Source

Facts

Specs and prices

FactDeepSeek-V4.1-Flash
MakerDeepSeek
API model iddeepseek-flash
ReleasedSeptember 10, 2026
StatusCurrent
Context window1M tokens
Max output384K tokens
Open weightsYes
Input, per 1M tokens$0.30
Cache hit, per 1M$0.006
Cache write, per 1M$0.30 (same as input)
Output, per 1M tokens$1.20
Runs inOpenCode and OpenRouter

Standard API rates in US dollars, as published by the maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. DeepSeek-V4.1-Flash: Prices are DeepSeek's peak rates. Off-peak hours cost 50% less: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays.

Good to know

  • Open weights under the MIT license.
  • The API id deepseek-flash serves it, and the older deepseek-v4-flash ids route to it.
  • The most-used model on OpenRouter in the week before September 28, 2026, by tokens processed.

Cost

What typical work costs

Example token counts at DeepSeek-V4.1-Flash’s published rates. On the agentic session, caching saves $0.59 against billing every token as ordinary input.

Example workload costs for DeepSeek-V4.1-Flash
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$0.22
Large one-off review, 150K input with no cache hits, 10K output$0.06
Output-heavy generation, 30K input, 80K output$0.11
A month of sessions, 110 sessions: 5 a day, 22 working days$24.42

Prompt caching

How DeepSeek bills cached tokens

DeepSeek

DeepSeek's disk cache is on by default for every account, with no code changes. DeepSeek lists no fee for writing the cache.

A cache hit costs $0.006 per million tokens on DeepSeek-V4.1-Flash and $0.044 on DeepSeek-V4-Pro at peak rates, and off-peak hours cost 50% less.

Source: DeepSeek API: Context caching

DeepSeek-V4.1-Flash compared

  • DeepSeek-V4.1-Flash vs GPT-5.6 Luna

    DeepSeek-V4.1-Flash and GPT-5.6 Luna both price a cached coding session at $0.22, from opposite ends: Luna's cheaper input, DeepSeek's cheaper cache hits.

  • DeepSeek-V4.1-Flash vs Gemini 3.8 Flash

    DeepSeek-V4.1-Flash costs $0.22 for a cached coding session that costs $0.71 on Gemini 3.8 Flash, and Gemini's introductory rates end December 31, 2026.

  • DeepSeek-V4.1-Flash vs Claude Haiku 4.5

    DeepSeek-V4.1-Flash holds 5x the context of Claude Haiku 4.5 and costs $0.22 against $1.20 for a cached coding session. Haiku 4.5's successor is announced.

  • DeepSeek-V4.1-Flash vs Claude Sonnet 5

    A cached coding session costs $0.22 on DeepSeek-V4.1-Flash and $2.40 on Claude Sonnet 5. Where the 10.9x gap comes from, and what open weights change.

  • DeepSeek-V4.1-Flash vs DeepSeek-V4-Pro

    DeepSeek says DeepSeek-V4.1-Flash outperforms its own DeepSeek-V4-Pro, which costs 4.3x as much per coding session. What Pro still offers, and what comes next.

  • DeepSeek-V4.1-Flash vs GLM-5.3-Flash

    GLM-5.3-Flash lists half DeepSeek-V4.1-Flash's input price, but GLM's cache hits cost 5x as much. A cached coding session: $0.16 against $0.22 at peak.

  • DeepSeek-V4.1-Flash vs GPT-6 Luna

    GPT-6 Luna undercuts DeepSeek-V4.1-Flash at list prices, $0.11 against $0.22 for a cached coding session. How cache pricing and off-peak hours move the gap.

  • DeepSeek-V4.1-Flash vs MiniMax M3

    DeepSeek-V4.1-Flash and MiniMax M3 list the same $0.30 input and $1.20 output rates, but MiniMax's cache hits cost 10x as much: $0.33 vs $0.22 a session.

Your own numbers

See what DeepSeek-V4.1-Flash really costs you.

everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math