Model comparison
DeepSeek-V4.1-Flash vs MiniMax M3: same rates, cheaper hits
DeepSeek-V4.1-Flash and MiniMax M3 list the same $0.30 input and $1.20 output rates, but MiniMax's cache hits cost 10x as much: $0.33 vs $0.22 a session.
· Prices as of September 28, 2026
DeepSeek-V4.1-Flash
DeepSeek · Released September 10, 2026
DeepSeek's small, low-cost, natively multimodal model on its new architecture, which DeepSeek says outperforms its own DeepSeek-V4-Pro.
DeepSeek-V4.1-Flash facts and comparisonsMiniMax M3
MiniMax · Released June 1, 2026
MiniMax's open-weight model pitched against closed frontier models, with a 1M context and native multimodality.
MiniMax M3 facts and comparisons
The short answer
DeepSeek-V4.1-Flash and MiniMax M3 list the same $0.30 input and $1.20 output per million, so uncached work costs the same on both, but a DeepSeek cache hit costs $0.006 against $0.06 and the example agentic coding session comes to $0.22 against $0.33. MiniMax M3 suits work with video input and lists a higher output limit, 524.3K tokens, though MiniMax recommends up to 131,072 per request, while DeepSeek-V4.1-Flash suits cache-heavy agents and off-peak work at 50% less. Both are open-weight models you reach through OpenRouter or OpenCode.
Choose DeepSeek-V4.1-Flash if
- Your agent resends a big cached context, and DeepSeek's hit costs $0.006 per million, a tenth of MiniMax's $0.06.
- You can run work off-peak, when every DeepSeek rate drops by 50%.
- You want the permissive MIT license rather than MiniMax's own community license.
- Your code already uses deepseek-flash or the older deepseek-v4-flash ids, which now route to V4.1 Flash.
Choose MiniMax M3 if
- You need responses longer than 384K tokens: MiniMax lists 524.3K, though it recommends up to 131,072 per request.
- You want video as well as image input, or desktop computer use, both of which MiniMax lists among M3's strengths.
- Your jobs are mostly uncached and run in DeepSeek's peak hours, where both cost the same per token and MiniMax's larger output limit comes at no extra charge.
Side by side
Specs and prices
| Fact | DeepSeek-V4.1-Flash | MiniMax M3 |
|---|---|---|
| Maker | DeepSeek | MiniMax |
| API model id | deepseek-flash | MiniMax-M3 |
| Released | September 10, 2026 | June 1, 2026 |
| Status | Current | Current |
| Context window | 1M tokens | 1M tokens |
| Max output | 384K tokens | 524.3K tokens |
| Open weights | Yes | Yes |
| Input, per 1M tokens | $0.30 | $0.30 |
| Cache hit, per 1M | $0.006 | $0.06 |
| Cache write, per 1M | $0.30 (same as input) | $0.30 (same as input) |
| Output, per 1M tokens | $1.20 | $1.20 |
| Runs in | OpenCode and OpenRouter | OpenCode and OpenRouter |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. DeepSeek-V4.1-Flash: Prices are DeepSeek's peak rates. Off-peak hours cost 50% less: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. MiniMax M3: MiniMax labels these rates a permanent 50% discount. Requests over 512K input tokens cost $0.60 input, $0.12 cached, and $2.40 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | DeepSeek-V4.1-Flash | MiniMax M3 |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.22 | $0.33 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.06 | $0.06 |
| Output-heavy generation, 30K input, 80K output | $0.11 | $0.11 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $24.42 | $36.30 |
| Where the session’s cost goes | ||
| Cache writes | $0.12 | $0.12 |
| Cache reads | $0.01 | $0.12 |
| Uncached input | $0.03 | $0.03 |
| Output | $0.06 | $0.06 |
- caching saves on the session with DeepSeek-V4.1-Flash (73%)
- $0.59
- caching saves on the session with MiniMax M3 (59%)
- $0.48
Identical rate cards, one line apart
Put the two price lists side by side and three of four lines match. DeepSeek-V4.1-Flash and MiniMax M3 both charge $0.30 per million input tokens and $1.20 per million output tokens, and neither charges to write the cache, so written tokens cost $0.30 as well. The large one-off review costs $0.06 on both and the output-heavy generation $0.11 on both.
The fourth line is the cache hit. MiniMax charges $0.06 per million, one-fifth of input. DeepSeek charges $0.006, 2% of input, which DeepSeek links to a KV cache that needs a quarter of the memory of the previous generation. The example session reads 2M tokens from the cache: $0.12 on MiniMax, 36% of its session, and $0.01 on DeepSeek.
So the session costs $0.33 on MiniMax M3 and $0.22 on DeepSeek-V4.1-Flash, 1.5x, and 110 sessions a month $36.30 against $24.42. Caching saves 73% of the uncached cost on DeepSeek and 59% on MiniMax. The more of each request a coding agent resends from the cache, the more that one line decides.
Two kinds of 50% discount: off-peak hours and a permanent cut
Both makers advertise a 50% cut, and they mean different things. MiniMax labels its current rates a permanent 50% discount, so the $0.30 and $1.20 shown here already include it. DeepSeek's $0.30 and $1.20 are peak rates, and DeepSeek charges 50% less outside peak.
DeepSeek's peak runs from 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. Off-peak, every DeepSeek rate halves, so uncached work that ties at peak costs half as much on DeepSeek, and the cached session gap grows.
MiniMax has a size rule instead. Requests over 512K input tokens cost $0.60 input, $0.12 cached, and $2.40 output per million, and the example session stays far below that line. MiniMax also caches only requests of 512 or more input tokens.
Output limits, inputs, and licenses
Both accept 1M tokens of context, and MiniMax credits its window to MiniMax Sparse Attention. Output limits differ: DeepSeek lists 384K per response, while MiniMax lists 524.3K and recommends up to 131,072 output tokens per request.
On inputs, DeepSeek describes V4.1 Flash as having native visual understanding. MiniMax lists image and video input and operating a desktop computer among M3's strengths, and says "M3 reaches frontier-level performance on specialized tasks such as coding and agentic work." DeepSeek, for its part, says tests by several parties put V4.1 Flash ahead of its own DeepSeek-V4-Pro.
The licenses differ. DeepSeek publishes V4.1 Flash under the MIT license, and MiniMax publishes M3 under the MiniMax community license, so read MiniMax's terms before building on its weights. Either set of weights can be self-hosted, at a hardware cost that falls outside this page's figures.
You reach DeepSeek-V4.1-Flash and MiniMax M3 through OpenRouter or OpenCode; Cursor and GitHub Copilot list neither. On OpenRouter, each request lands with one of several providers, and those providers can charge differently from the makers' own APIs, whose prices these tables use.
Prompt caching
How each maker bills cached tokens
DeepSeek
DeepSeek's disk cache is on by default for every account, with no code changes. DeepSeek lists no fee for writing the cache.
A cache hit costs $0.006 per million tokens on DeepSeek-V4.1-Flash and $0.044 on DeepSeek-V4-Pro at peak rates, and off-peak hours cost 50% less.
Source: DeepSeek API: Context caching
MiniMax
MiniMax M3 caches repeated prompts automatically for requests of 512 or more input tokens, with no charge for cache writes.
A cache hit costs $0.06 per million tokens, one-fifth of the input price.
Source: MiniMax docs: Prompt caching
Your own numbers
See what DeepSeek-V4.1-Flash and MiniMax M3 really cost you.
everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Do DeepSeek-V4.1-Flash and MiniMax M3 cost the same?
Per input and output token, yes: $0.30 and $1.20 per million on both, at DeepSeek's peak rates. Cache hits differ, $0.006 against $0.06, so the example cached session costs $0.22 on DeepSeek and $0.33 on MiniMax.
What is MiniMax's permanent 50% discount?
MiniMax labels its current M3 rates a permanent 50% discount, and those discounted rates are the ones shown here. Requests over 512K input tokens cost $0.60 input, $0.12 cached, and $2.40 output per million.
Which model can write longer responses?
MiniMax M3 lists 524.3K output tokens and recommends up to 131,072 per request. DeepSeek-V4.1-Flash lists 384K.
How can I track what each one costs me?
EveryToken prices DeepSeek-V4.1-Flash and MiniMax M3 from OpenRouter's catalog whenever you use them through OpenRouter. It does not price calls to DeepSeek's or MiniMax's own APIs.