DeepSeek
DeepSeek-V4.1-Flash: price, context window, and caching
DeepSeek's small, low-cost, natively multimodal model on its new architecture, which DeepSeek says outperforms its own DeepSeek-V4-Pro.
Released September 10, 2026 · Prices as of September 28, 2026
In DeepSeek’s words
“Introducing the smallest model in our new architecture family, with native visual understanding.”
Facts
Specs and prices
| Fact | DeepSeek-V4.1-Flash |
|---|---|
| Maker | DeepSeek |
| API model id | deepseek-flash |
| Released | September 10, 2026 |
| Status | Current |
| Context window | 1M tokens |
| Max output | 384K tokens |
| Open weights | Yes |
| Input, per 1M tokens | $0.30 |
| Cache hit, per 1M | $0.006 |
| Cache write, per 1M | $0.30 (same as input) |
| Output, per 1M tokens | $1.20 |
| Runs in | OpenCode and OpenRouter |
Standard API rates in US dollars, as published by the maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. DeepSeek-V4.1-Flash: Prices are DeepSeek's peak rates. Off-peak hours cost 50% less: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays.
Good to know
- Open weights under the MIT license.
- The API id deepseek-flash serves it, and the older deepseek-v4-flash ids route to it.
- The most-used model on OpenRouter in the week before September 28, 2026, by tokens processed.
Cost
What typical work costs
Example token counts at DeepSeek-V4.1-Flash’s published rates. On the agentic session, caching saves $0.59 against billing every token as ordinary input.
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.22 |
|---|---|
| Large one-off review, 150K input with no cache hits, 10K output | $0.06 |
| Output-heavy generation, 30K input, 80K output | $0.11 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $24.42 |
Prompt caching
How DeepSeek bills cached tokens
DeepSeek
DeepSeek's disk cache is on by default for every account, with no code changes. DeepSeek lists no fee for writing the cache.
A cache hit costs $0.006 per million tokens on DeepSeek-V4.1-Flash and $0.044 on DeepSeek-V4-Pro at peak rates, and off-peak hours cost 50% less.
Source: DeepSeek API: Context caching
DeepSeek-V4.1-Flash compared
DeepSeek-V4.1-Flash vs GPT-5.6 Luna
DeepSeek-V4.1-Flash and GPT-5.6 Luna both price a cached coding session at $0.22, from opposite ends: Luna's cheaper input, DeepSeek's cheaper cache hits.
DeepSeek-V4.1-Flash vs Gemini 3.8 Flash
DeepSeek-V4.1-Flash costs $0.22 for a cached coding session that costs $0.71 on Gemini 3.8 Flash, and Gemini's introductory rates end December 31, 2026.
DeepSeek-V4.1-Flash vs Claude Haiku 4.5
DeepSeek-V4.1-Flash holds 5x the context of Claude Haiku 4.5 and costs $0.22 against $1.20 for a cached coding session. Haiku 4.5's successor is announced.
DeepSeek-V4.1-Flash vs Claude Sonnet 5
A cached coding session costs $0.22 on DeepSeek-V4.1-Flash and $2.40 on Claude Sonnet 5. Where the 10.9x gap comes from, and what open weights change.
DeepSeek-V4.1-Flash vs DeepSeek-V4-Pro
DeepSeek says DeepSeek-V4.1-Flash outperforms its own DeepSeek-V4-Pro, which costs 4.3x as much per coding session. What Pro still offers, and what comes next.
DeepSeek-V4.1-Flash vs GLM-5.3-Flash
GLM-5.3-Flash lists half DeepSeek-V4.1-Flash's input price, but GLM's cache hits cost 5x as much. A cached coding session: $0.16 against $0.22 at peak.
DeepSeek-V4.1-Flash vs GPT-6 Luna
GPT-6 Luna undercuts DeepSeek-V4.1-Flash at list prices, $0.11 against $0.22 for a cached coding session. How cache pricing and off-peak hours move the gap.
DeepSeek-V4.1-Flash vs MiniMax M3
DeepSeek-V4.1-Flash and MiniMax M3 list the same $0.30 input and $1.20 output rates, but MiniMax's cache hits cost 10x as much: $0.33 vs $0.22 a session.
Your own numbers
See what DeepSeek-V4.1-Flash really costs you.
everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.