Model comparison
DeepSeek-V4.1-Flash and GPT-5.6 Luna tie on session cost
DeepSeek-V4.1-Flash and GPT-5.6 Luna both price a cached coding session at $0.22, from opposite ends: Luna's cheaper input, DeepSeek's cheaper cache hits.
· Prices as of September 28, 2026
DeepSeek-V4.1-Flash
DeepSeek · Released September 10, 2026
DeepSeek's small, low-cost, natively multimodal model on its new architecture, which DeepSeek says outperforms its own DeepSeek-V4-Pro.
DeepSeek-V4.1-Flash facts and comparisonsGPT-5.6 Luna
OpenAI · Released July 9, 2026 · Previous generation
The low-cost GPT-5.6 tier for cost-sensitive, high-volume work, which roughly corresponds to earlier GPT-5 nano models.
GPT-5.6 Luna facts and comparisons
The short answer
DeepSeek-V4.1-Flash and GPT-5.6 Luna cost the same $0.22 for the example agentic coding session, and 110 sessions a month differ by $0.22, because Luna's lower input price and DeepSeek's lower cache-hit price cancel out. GPT-5.6 Luna is a previous-generation model that Codex suggests replacing with GPT-6 Luna, so it suits teams already running it in Codex, Cursor, or GitHub Copilot. DeepSeek-V4.1-Flash is the current model of the two, with open weights and a 50% off-peak discount that tips the session its way.
Choose DeepSeek-V4.1-Flash if
- You can schedule work outside DeepSeek's peak hours, when it charges 50% less.
- Your agent resends a large cached context: a DeepSeek hit costs $0.006 per million against $0.02, under a third of the price.
- You want a current model with open weights under the MIT license rather than a previous-generation tier.
- You need responses longer than 128K tokens, since DeepSeek lists 384K.
Choose GPT-5.6 Luna if
- Your work is mostly uncached: Luna's $0.20 input makes the large one-off review $0.04 against $0.06.
- You already use it in Codex, Cursor, or GitHub Copilot and want to keep a working setup while you evaluate GPT-6 Luna.
- You want explicit cache breakpoints, up to four, which OpenAI supports from GPT-5.6 on.
- Your team works during DeepSeek's peak hours, when its full rates apply.
Side by side
Specs and prices
| Fact | DeepSeek-V4.1-Flash | GPT-5.6 Luna |
|---|---|---|
| Maker | DeepSeek | OpenAI |
| API model id | deepseek-flash | gpt-5.6-luna |
| Released | September 10, 2026 | July 9, 2026 |
| Status | Current | Previous generation |
| Context window | 1M tokens | 1.05M tokens |
| Max output | 384K tokens | 128K tokens |
| Open weights | Yes | No |
| Input, per 1M tokens | $0.30 | $0.20 |
| Cache hit, per 1M | $0.006 | $0.02 |
| Cache write, per 1M | $0.30 (same as input) | $0.25 |
| Output, per 1M tokens | $1.20 | $1.20 |
| Runs in | OpenCode and OpenRouter | Codex, Cursor, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. DeepSeek-V4.1-Flash: Prices are DeepSeek's peak rates. Off-peak hours cost 50% less: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. GPT-5.6 Luna: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | DeepSeek-V4.1-Flash | GPT-5.6 Luna |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.22 | $0.22 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.06 | $0.04 |
| Output-heavy generation, 30K input, 80K output | $0.11 | $0.10 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $24.42 | $24.20 |
| Where the session’s cost goes | ||
| Cache writes | $0.12 | $0.10 |
| Cache reads | $0.01 | $0.04 |
| Uncached input | $0.03 | $0.02 |
| Output | $0.06 | $0.06 |
- caching saves on the session with DeepSeek-V4.1-Flash (73%)
- $0.59
- caching saves on the session with GPT-5.6 Luna (61%)
- $0.34
How two different price lists reach the same $0.22
The rate cards look nothing alike. GPT-5.6 Luna charges $0.20 per million input tokens, $0.25 per million for a cache write, and $0.02 per million for a cache hit. DeepSeek-V4.1-Flash charges $0.30 for input, bills cache writes as ordinary input because DeepSeek lists no fee for writing its cache, and charges $0.006 for a hit. Output is $1.20 per million on both.
Line by line, the example session trades evenly. Luna saves $0.02 on cache writes, which are 45% of its session, and $0.01 on fresh input. DeepSeek saves $0.03 on cache reads, where 2M tokens cost $0.01 against $0.04. Output costs $0.06 on each. Both sessions round to $0.22, and 110 of them come to $24.42 on DeepSeek and $24.20 on Luna.
Uncached work tilts toward Luna. The large one-off review costs $0.04 against $0.06, a 1.5x gap that is just the input ratio, and the output-heavy generation costs $0.10 against $0.11. The more of its context an agent reads from the cache, the better DeepSeek's side looks; the more fresh input it sends, the better Luna's does.
GPT-5.6 Luna is now a previous-generation model
OpenAI released GPT-5.6 Luna on July 9, 2026 as the low-cost GPT-5.6 tier, which roughly corresponds to earlier GPT-5 nano models, and cut its price by 80% on July 30, 2026. It is now a previous generation. Codex suggests moving to GPT-6 Luna, and GPT-5.6 Luna stays available in the API.
That matters for a long-term choice. OpenAI describes GPT-5.6 Luna as a "GPT-5.6 model optimized for cost-sensitive workloads" and says the GPT-5.6 family reaches flagship-level performance with fewer output tokens. Its successor, GPT-6 Luna, lists lower rates, and Codex points new work there.
DeepSeek-V4.1-Flash is the newer model, released on September 10, 2026. DeepSeek introduces it as "the smallest model in our new architecture family, with native visual understanding" and says tests by several parties put it ahead of DeepSeek-V4-Pro on performance, cost, speed, and total runtime.
Off-peak hours, OpenRouter, and open weights
The DeepSeek rates in these tables are peak rates, which apply from 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. At all other times DeepSeek charges 50% less. With the session tied at peak, a session run off-peak costs about half as much on DeepSeek as on GPT-5.6 Luna.
OpenRouter lists both models. For an open-weight model like DeepSeek-V4.1-Flash, it sends the request to one of several providers, whose prices can differ from DeepSeek's own; the tables use each maker's own price. GPT-5.6 Luna is also in Codex, Cursor, OpenCode, and GitHub Copilot, while DeepSeek-V4.1-Flash is in OpenCode but not in Cursor or Copilot.
DeepSeek publishes the model's weights under the MIT license, so you can also run it yourself. This page does not estimate hosting costs, which depend on hardware and usage. EveryToken prices GPT-5.6 Luna from your Codex, Cursor, or OpenCode history at OpenAI's rates, and prices DeepSeek-V4.1-Flash when you use it through OpenRouter.
Prompt caching
How each maker bills cached tokens
DeepSeek
DeepSeek's disk cache is on by default for every account, with no code changes. DeepSeek lists no fee for writing the cache.
A cache hit costs $0.006 per million tokens on DeepSeek-V4.1-Flash and $0.044 on DeepSeek-V4-Pro at peak rates, and off-peak hours cost 50% less.
Source: DeepSeek API: Context caching
OpenAI
Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.
From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.
On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.
Source: OpenAI: Prompt caching
Your own numbers
See what DeepSeek-V4.1-Flash and GPT-5.6 Luna really cost you.
everyaitoken reads your OpenRouter, Codex, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Do DeepSeek-V4.1-Flash and GPT-5.6 Luna really cost the same?
On the example cached session, yes: both come to $0.22 at published rates, DeepSeek's at its peak. They differ on other work: a large uncached review costs $0.06 on DeepSeek and $0.04 on Luna. Off-peak, DeepSeek charges 50% less.
Should I move from GPT-5.6 Luna to GPT-6 Luna?
Codex suggests that move, and GPT-5.6 Luna stays available in the API meanwhile. The GPT-6 Luna model page lists its rates and limits for comparison.
Which model has cheaper cache hits?
DeepSeek-V4.1-Flash: $0.006 per million, 2% of its input price, against $0.02 on GPT-5.6 Luna, 10% of input. OpenAI also charges 1.25x input to write the cache from GPT-5.6 on, while DeepSeek lists no write fee.
What is the output limit on each?
DeepSeek-V4.1-Flash lists 384K output tokens and GPT-5.6 Luna 128K. Their context windows are 1M and 1.05M tokens.
Sources
- DeepSeek API: Models and pricing
- DeepSeek: V4.1 Flash release
- DeepSeek API: Change log
- OpenRouter: DeepSeek-V4.1-Flash
- OpenRouter: Rankings
- OpenCode docs: Zen
- OpenAI: API pricing
- OpenAI docs: GPT-5.6 Luna
- OpenAI: Using GPT-5.6
- OpenAI: API changelog
- Codex docs: Models
- Cursor docs: Models and pricing
- OpenRouter: GPT-5.6 Luna
- DeepSeek API: Context caching
- OpenAI: Prompt caching