Model comparison
DeepSeek-V4.1-Flash vs GPT-6 Luna: which costs less?
GPT-6 Luna undercuts DeepSeek-V4.1-Flash at list prices, $0.11 against $0.22 for a cached coding session. How cache pricing and off-peak hours move the gap.
· Prices as of September 28, 2026
DeepSeek-V4.1-Flash
DeepSeek · Released September 10, 2026
DeepSeek's small, low-cost, natively multimodal model on its new architecture, which DeepSeek says outperforms its own DeepSeek-V4-Pro.
DeepSeek-V4.1-Flash facts and comparisonsGPT-6 Luna
OpenAI · Released September 22, 2026
The cheapest GPT-6 model, pitched for focused, high-volume, repeatable work, including narrower coding tasks.
GPT-6 Luna facts and comparisons
The short answer
GPT-6 Luna is the cheaper of the two at published rates, $0.11 against $0.22 for the example agentic coding session, because its $0.10 input and $0.50 output undercut DeepSeek-V4.1-Flash on every uncached token. DeepSeek-V4.1-Flash wins back ground on cache hits and costs 50% less off-peak, which brings the session close to even. Choose on where you work: GPT-6 Luna runs in Codex, while DeepSeek-V4.1-Flash is an open-weight model you reach through OpenRouter, OpenCode, or your own servers.
Choose DeepSeek-V4.1-Flash if
- You want open weights under the MIT license, so the model can run on hardware you control.
- Your work runs mostly outside 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, when DeepSeek charges 50% less.
- Your agent leans on the cache: a DeepSeek hit costs $0.006 per million against $0.01 on GPT-6 Luna, and DeepSeek lists no fee for writing the cache.
- You need up to 384K tokens of output in one response, 3x the 128K GPT-6 Luna allows.
Choose GPT-6 Luna if
- You want the lower price on uncached work: $0.10 input and $0.50 output per million, against $0.30 and $1.20 on DeepSeek at peak.
- You work in Codex, whose docs recommend GPT-6 Luna for focused, repeatable tasks and which supports effort up to max for it.
- You use GitHub Copilot, or the Codex app on a Free or Go plan, both of which offer GPT-6 Luna.
- Your team works during DeepSeek's peak hours, when its full list rates apply.
Side by side
Specs and prices
| Fact | DeepSeek-V4.1-Flash | GPT-6 Luna |
|---|---|---|
| Maker | DeepSeek | OpenAI |
| API model id | deepseek-flash | gpt-6-luna |
| Released | September 10, 2026 | September 22, 2026 |
| Status | Current | Current |
| Context window | 1M tokens | 1.05M tokens |
| Max output | 384K tokens | 128K tokens |
| Open weights | Yes | No |
| Input, per 1M tokens | $0.30 | $0.10 |
| Cache hit, per 1M | $0.006 | $0.01 |
| Cache write, per 1M | $0.30 (same as input) | $0.125 |
| Output, per 1M tokens | $1.20 | $0.50 |
| Runs in | OpenCode and OpenRouter | Codex, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. DeepSeek-V4.1-Flash: Prices are DeepSeek's peak rates. Off-peak hours cost 50% less: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. GPT-6 Luna: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | DeepSeek-V4.1-Flash | GPT-6 Luna |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.22 | $0.11 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.06 | $0.02 |
| Output-heavy generation, 30K input, 80K output | $0.11 | $0.04 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $24.42 | $11.55 |
| Where the session’s cost goes | ||
| Cache writes | $0.12 | $0.05 |
| Cache reads | $0.01 | $0.02 |
| Uncached input | $0.03 | $0.01 |
| Output | $0.06 | $0.03 |
- caching saves on the session with DeepSeek-V4.1-Flash (73%)
- $0.59
- caching saves on the session with GPT-6 Luna (61%)
- $0.17
Why GPT-6 Luna comes out cheaper at list prices
GPT-6 Luna is the cheapest model in OpenAI's GPT-6 family: $0.10 per million input tokens, $0.50 per million output tokens, and $0.01 per million cache hits. DeepSeek-V4.1-Flash lists $0.30 input and $1.20 output at DeepSeek's peak rates. On uncached work that is a 3x gap on the large one-off review, $0.06 against $0.02, and 2.8x on the output-heavy generation, $0.11 against $0.04.
The cache narrows it but does not close it. DeepSeek's cache hit is $0.006 per million, 2% of its input price, and OpenAI's is $0.01, 10% of Luna's. DeepSeek lists no fee for writing the cache, while OpenAI charges 1.25x input for a write from GPT-5.6 on, $0.125 on Luna. Because Luna's input is so low, that premium still leaves its writes at $0.05 for the session's 400K written tokens, against $0.12 on DeepSeek.
The session lands at $0.22 on DeepSeek-V4.1-Flash and $0.11 on GPT-6 Luna, a 2x gap, and 110 sessions a month come to $24.42 against $11.55. Cache writes are the largest line on both, 55% of DeepSeek's session and 45% of Luna's, so the write price matters more here than the headline input rate.
Does DeepSeek's off-peak discount change the answer?
It narrows it a lot. The DeepSeek prices on this page are peak rates, and DeepSeek charges 50% less off-peak, cache hits included. Peak runs from 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays, which is seven hours of a weekday and none of the weekend.
Halving every DeepSeek line brings the example session close to what GPT-6 Luna charges, though Luna stays slightly lower. Uncached work still favors Luna, because DeepSeek's discounted input and output rates remain above Luna's $0.10 and $0.50.
The discount is listed on DeepSeek's own price page, so it applies to DeepSeek's API. OpenRouter sends a request for an open-weight model to one of several providers, and each sets its own price, which can differ from DeepSeek's. The tables here use each maker's own price.
Context, output, and long prompts
GPT-6 Luna has a 1.05M context window and DeepSeek-V4.1-Flash 1M, so both can hold a large repository. Luna carries a long-context rule: once a request passes 272K input tokens, OpenAI bills all of it at 2x for input and cache and 1.5x for output. Every request in the example session stays below 200K, so the Luna figures above use its standard rates.
Output limits differ more. DeepSeek lists 384K output tokens against 128K for Luna. OpenAI sets Luna's reasoning effort at medium by default and lets Codex raise it to max, and more effort means more output tokens. Output is 27% of the session on both models here, so effort settings move both totals.
OpenAI calls GPT-6 Luna "our most efficient model for focused, high-volume tasks," and the Codex docs recommend it for focused, repeatable work. DeepSeek presents V4.1 Flash as the smallest model in its new architecture family, with native visual understanding, and says its KV cache needs a quarter of the memory of the previous generation, cutting cache-hit costs for agents.
EveryToken prices GPT-6 Luna at OpenAI's rates from your Codex or OpenCode history, and prices DeepSeek-V4.1-Flash only when you use it through OpenRouter, from OpenRouter's catalog.
Prompt caching
How each maker bills cached tokens
DeepSeek
DeepSeek's disk cache is on by default for every account, with no code changes. DeepSeek lists no fee for writing the cache.
A cache hit costs $0.006 per million tokens on DeepSeek-V4.1-Flash and $0.044 on DeepSeek-V4-Pro at peak rates, and off-peak hours cost 50% less.
Source: DeepSeek API: Context caching
OpenAI
Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.
From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.
On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.
Source: OpenAI: Prompt caching
Your own numbers
See what DeepSeek-V4.1-Flash and GPT-6 Luna really cost you.
everyaitoken reads your OpenRouter, Codex, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is GPT-6 Luna cheaper than DeepSeek-V4.1-Flash?
At published rates, yes. Input costs $0.10 against $0.30, output $0.50 against $1.20, and the example cached session $0.11 against $0.22. DeepSeek's cache hits are cheaper, $0.006 against $0.01 per million, but not by enough to make up the difference.
When is DeepSeek-V4.1-Flash off-peak?
Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. Every other hour is off-peak and costs 50% less, including cache hits.
Where can I use each model?
GPT-6 Luna runs in Codex, OpenRouter, OpenCode, and GitHub Copilot, but not in Codex cloud. DeepSeek-V4.1-Flash is available through OpenRouter and OpenCode and on DeepSeek's own API as deepseek-flash, and it is not offered in Cursor or Copilot.
Does GPT-6 Luna have a long-context surcharge?
Yes. A Luna request that passes 272K input tokens is billed in full at 2x for input and cache and 1.5x for output, not just the part above the line. Below it, the $0.10 input and $0.50 output rates apply.
Sources
- DeepSeek API: Models and pricing
- DeepSeek: V4.1 Flash release
- DeepSeek API: Change log
- OpenRouter: DeepSeek-V4.1-Flash
- OpenRouter: Rankings
- OpenCode docs: Zen
- OpenAI: API pricing
- OpenAI docs: GPT-6 Luna
- OpenAI: Introducing GPT-6 Sol and Luna
- OpenAI: API changelog
- Codex docs: Models
- OpenRouter: GPT-6 Luna
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- DeepSeek API: Context caching
- OpenAI: Prompt caching