Model comparison
DeepSeek-V4.1-Flash vs Claude Sonnet 5: a 10.9x cost gap
A cached coding session costs $0.22 on DeepSeek-V4.1-Flash and $2.40 on Claude Sonnet 5. Where the 10.9x gap comes from, and what open weights change.
· Prices as of September 28, 2026
DeepSeek-V4.1-Flash
DeepSeek · Released September 10, 2026
DeepSeek's small, low-cost, natively multimodal model on its new architecture, which DeepSeek says outperforms its own DeepSeek-V4-Pro.
DeepSeek-V4.1-Flash facts and comparisonsClaude Sonnet 5
Anthropic · Released June 30, 2026
Anthropic's balance of speed and intelligence, and a drop-in upgrade for Claude Sonnet 4.6. Anthropic pitches it as close to Claude Opus 4.8 at a lower price.
Claude Sonnet 5 facts and comparisons
The short answer
DeepSeek-V4.1-Flash prices the example agentic coding session at $0.22 against $2.40 on Claude Sonnet 5, a 10.9x gap that is wider than the 6.7x gap on the large uncached review because DeepSeek's cache hits cost $0.006 per million and writing the cache carries no premium. Pick Claude Sonnet 5 if you work in Claude Code or want what Anthropic calls its most agentic Sonnet yet, at its published rates. Pick DeepSeek-V4.1-Flash if cost per session drives the choice and you can run it through OpenRouter, OpenCode, or your own hardware, since its weights are open under the MIT license.
Choose DeepSeek-V4.1-Flash if
- Cost per session decides it: the example session costs $0.22 here and $2.40 on Claude Sonnet 5, and 110 sessions come to $24.42 against $264.00.
- Your agent rereads a large cached context, since a DeepSeek cache hit costs $0.006 per million, 2% of its input price, with no fee for writing the cache.
- You want open weights under the MIT license, so the model can run on infrastructure you control.
- You need long single responses: DeepSeek lists up to 384K output tokens, against 128K on Sonnet 5.
Choose Claude Sonnet 5 if
- You work in Claude Code, Anthropic's own agent, where the sonnet alias resolves to Claude Sonnet 5 on the Anthropic API.
- You want the model Anthropic describes as close to Claude Opus 4.8 at lower prices, built for planning and for using tools like browsers and terminals on its own.
- Your team also works in Cursor or GitHub Copilot, which offer Sonnet 5 and do not list DeepSeek-V4.1-Flash.
Side by side
Specs and prices
| Fact | DeepSeek-V4.1-Flash | Claude Sonnet 5 |
|---|---|---|
| Maker | DeepSeek | Anthropic |
| API model id | deepseek-flash | claude-sonnet-5 |
| Released | September 10, 2026 | June 30, 2026 |
| Status | Current | Current |
| Context window | 1M tokens | 1M tokens |
| Max output | 384K tokens | 128K tokens |
| Open weights | Yes | No |
| Input, per 1M tokens | $0.30 | $2 |
| Cache hit, per 1M | $0.006 | $0.20 |
| Cache write, per 1M | $0.30 (same as input) | $2.50 (5-minute), $4 (1-hour) |
| Output, per 1M tokens | $1.20 | $10 |
| Runs in | OpenCode and OpenRouter | Claude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker (DeepSeek-V4.1-Flash: September 28, 2026; Claude Sonnet 5: September 26, 2026). Batch and priority tiers, taxes, and subscription plans are not included. DeepSeek-V4.1-Flash: Prices are DeepSeek's peak rates. Off-peak hours cost 50% less: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. Claude Sonnet 5: The full 1M context window is billed at standard rates. The launch price became the standard price on August 10, 2026, and a planned increase was cancelled.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | DeepSeek-V4.1-Flash | Claude Sonnet 5 |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.22 | $2.40 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.06 | $0.40 |
| Output-heavy generation, 30K input, 80K output | $0.11 | $0.86 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $24.42 | $264.00 |
| Where the session’s cost goes | ||
| Cache writes | $0.12 | $1.30 |
| Cache reads | $0.01 | $0.40 |
| Uncached input | $0.03 | $0.20 |
| Output | $0.06 | $0.50 |
- caching saves on the session with DeepSeek-V4.1-Flash (73%)
- $0.59
- caching saves on the session with Claude Sonnet 5 (56%)
- $3.10
Why the gap is wider on a cached session than on one-off work
Start with the rate cards. DeepSeek-V4.1-Flash charges $0.30 per million input tokens and $1.20 per million output tokens at DeepSeek's peak rates. Claude Sonnet 5 charges $2 and $10. That is 6.7x on input and 8.3x on output, and the large one-off review, which has no cache hits, lands at the input ratio: $0.06 against $0.40.
Caching pulls the two further apart, not closer. A DeepSeek cache hit costs $0.006 per million, 2% of the input price, while a Sonnet 5 hit costs $0.20, the usual 10%. That is a 33.3x difference on the line that carries most of the tokens for any agent that resends its context. In the example session, 2M cached tokens cost $0.01 on DeepSeek and $0.40 on Sonnet 5.
Cache writes add to it. DeepSeek lists no fee for writing the cache, so the 400K written tokens cost ordinary input, $0.12. Anthropic bills a 5-minute write at 1.25x input and a 1-hour write at 2x, which is $2.50 and $4 per million on Sonnet 5, so the same writes cost $1.30. Together the session costs $0.22 on DeepSeek and $2.40 on Sonnet 5, a 10.9x gap, against 7.8x on the output-heavy generation.
What off-peak pricing, OpenRouter, and open weights do to these numbers
Every DeepSeek figure on this page is a peak rate. Peak covers 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays, and DeepSeek charges 50% less at all other times. That leaves most of the working day in the Americas and every weekend at the lower rate, so the gap to Sonnet 5 can be twice as wide as the tables show.
The tables also use each maker's own API price. Through OpenRouter, a request for an open-weight model such as DeepSeek-V4.1-Flash goes to one of several providers, whose prices can differ from DeepSeek's own. It was the most-used model on OpenRouter in the week before September 28, 2026, by tokens processed. Claude Sonnet 5 is on OpenRouter too, but the Claude figures here are Anthropic's rates.
Open weights change the options, not the table. DeepSeek publishes V4.1 Flash under the MIT license, so a team can run it on its own hardware or a cloud of its choice. What that costs depends on the hardware and how busy it stays, and this page does not estimate it. Claude Sonnet 5 has no open weights: you reach it through Anthropic's API and the tools that offer it.
Claude Code or OpenCode: where each model runs
Tool choice may settle this before price does. Claude Code is Anthropic's own coding agent and runs Claude models, with the sonnet alias pointing to Sonnet 5 on the Anthropic API. Sonnet 5 is also offered in Cursor, OpenRouter, OpenCode, and GitHub Copilot. DeepSeek-V4.1-Flash is not in Cursor or Copilot; you reach it through OpenRouter, OpenCode, DeepSeek's own API under the id deepseek-flash, or your own deployment.
The makers pitch the models differently. Anthropic says "Claude Sonnet 5 is built to be the most agentic Sonnet model yet" and describes performance close to Claude Opus 4.8 at lower prices. DeepSeek introduces V4.1 Flash as "the smallest model in our new architecture family, with native visual understanding," and says tests by several parties put it ahead of its own DeepSeek-V4-Pro on performance, cost, speed, and total runtime. Neither claim compares the two models directly, and this page makes no such comparison either.
Token counts do not line up one to one. Sonnet 5's newer tokenizer counts about 1.0 to 1.35x as many tokens as Claude Sonnet 4.6 for the same text, and the two makers' tokenizers differ, so the same repository comes out at different sizes. EveryToken prices Sonnet 5 at Anthropic's rates from your Claude Code history, and prices DeepSeek-V4.1-Flash when you run it through OpenRouter, from OpenRouter's catalog, so you can compare the two on your own sessions.
Prompt caching
How each maker bills cached tokens
DeepSeek
DeepSeek's disk cache is on by default for every account, with no code changes. DeepSeek lists no fee for writing the cache.
A cache hit costs $0.006 per million tokens on DeepSeek-V4.1-Flash and $0.044 on DeepSeek-V4-Pro at peak rates, and off-peak hours cost 50% less.
Source: DeepSeek API: Context caching
Anthropic
Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.
A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.
Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.
In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.
Source: Anthropic: Prompt caching
Your own numbers
See what DeepSeek-V4.1-Flash and Claude Sonnet 5 really cost you.
everyaitoken reads your OpenRouter, Claude Code, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
How much cheaper is DeepSeek-V4.1-Flash than Claude Sonnet 5?
At list prices, 6.7x on input and 8.3x on output. On the example cached session the gap grows to 10.9x, $0.22 against $2.40, because DeepSeek prices cache hits and cache writes far lower. Over 110 sessions a month that is $24.42 against $264.00, at DeepSeek's peak rates.
Can I use DeepSeek-V4.1-Flash in Claude Code?
Claude Code is Anthropic's own agent, built around Claude models and defaulting to them. Pointing it at another provider's compatible endpoint takes custom configuration that this comparison leaves out. DeepSeek-V4.1-Flash runs in OpenCode and through OpenRouter, and DeepSeek serves it on its own API as deepseek-flash, with the older deepseek-v4-flash ids routing to it.
Does DeepSeek charge for cache writes?
DeepSeek lists no fee for writing the cache, and its disk cache is on by default for every account with no code changes. Written tokens are billed as ordinary input in the tables here. Anthropic charges 1.25x input for a 5-minute write and 2x for a 1-hour write.
Is DeepSeek-V4.1-Flash open source?
Its weights are open under the MIT license, so you can download and self-host it. Hosting costs depend on your hardware and are not part of the figures on this page, which use DeepSeek's own API prices.
Do the two models have the same context window?
Yes, both accept 1M tokens. DeepSeek lists up to 384K output tokens against 128K for Claude Sonnet 5, which matters only when one response has to be very long.
Sources
- DeepSeek API: Models and pricing
- DeepSeek: V4.1 Flash release
- DeepSeek API: Change log
- OpenRouter: DeepSeek-V4.1-Flash
- OpenRouter: Rankings
- OpenCode docs: Zen
- Anthropic: Pricing
- Anthropic docs: Claude Sonnet 5
- Anthropic: Introducing Claude Sonnet 5
- Claude Code docs: Model configuration
- Cursor docs: Claude Sonnet 5
- OpenRouter: Claude Sonnet 5
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- DeepSeek API: Context caching
- Anthropic: Prompt caching