Skip to content

Model comparison

DeepSeek-V4.1-Flash vs GPT-6 Luna: which costs less?

GPT-6 Luna undercuts DeepSeek-V4.1-Flash at list prices, $0.11 against $0.22 for a cached coding session. How cache pricing and off-peak hours move the gap.

· Prices as of September 28, 2026

  • DeepSeek-V4.1-Flash

    DeepSeek · Released September 10, 2026

    DeepSeek's small, low-cost, natively multimodal model on its new architecture, which DeepSeek says outperforms its own DeepSeek-V4-Pro.

    DeepSeek-V4.1-Flash facts and comparisons
  • GPT-6 Luna

    OpenAI · Released September 22, 2026

    The cheapest GPT-6 model, pitched for focused, high-volume, repeatable work, including narrower coding tasks.

    GPT-6 Luna facts and comparisons

The short answer

GPT-6 Luna is the cheaper of the two at published rates, $0.11 against $0.22 for the example agentic coding session, because its $0.10 input and $0.50 output undercut DeepSeek-V4.1-Flash on every uncached token. DeepSeek-V4.1-Flash wins back ground on cache hits and costs 50% less off-peak, which brings the session close to even. Choose on where you work: GPT-6 Luna runs in Codex, while DeepSeek-V4.1-Flash is an open-weight model you reach through OpenRouter, OpenCode, or your own servers.

Choose DeepSeek-V4.1-Flash if

  • You want open weights under the MIT license, so the model can run on hardware you control.
  • Your work runs mostly outside 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, when DeepSeek charges 50% less.
  • Your agent leans on the cache: a DeepSeek hit costs $0.006 per million against $0.01 on GPT-6 Luna, and DeepSeek lists no fee for writing the cache.
  • You need up to 384K tokens of output in one response, 3x the 128K GPT-6 Luna allows.

Choose GPT-6 Luna if

  • You want the lower price on uncached work: $0.10 input and $0.50 output per million, against $0.30 and $1.20 on DeepSeek at peak.
  • You work in Codex, whose docs recommend GPT-6 Luna for focused, repeatable tasks and which supports effort up to max for it.
  • You use GitHub Copilot, or the Codex app on a Free or Go plan, both of which offer GPT-6 Luna.
  • Your team works during DeepSeek's peak hours, when its full list rates apply.

Side by side

Specs and prices

FactDeepSeek-V4.1-FlashGPT-6 Luna
MakerDeepSeekOpenAI
API model iddeepseek-flashgpt-6-luna
ReleasedSeptember 10, 2026September 22, 2026
StatusCurrentCurrent
Context window1M tokens1.05M tokens
Max output384K tokens128K tokens
Open weightsYesNo
Input, per 1M tokens$0.30$0.10
Cache hit, per 1M$0.006$0.01
Cache write, per 1M$0.30 (same as input)$0.125
Output, per 1M tokens$1.20$0.50
Runs inOpenCode and OpenRouterCodex, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. DeepSeek-V4.1-Flash: Prices are DeepSeek's peak rates. Off-peak hours cost 50% less: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. GPT-6 Luna: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadDeepSeek-V4.1-FlashGPT-6 Luna
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$0.22$0.11
Large one-off review, 150K input with no cache hits, 10K output$0.06$0.02
Output-heavy generation, 30K input, 80K output$0.11$0.04
A month of sessions, 110 sessions: 5 a day, 22 working days$24.42$11.55
Where the session’s cost goes
Cache writes$0.12$0.05
Cache reads$0.01$0.02
Uncached input$0.03$0.01
Output$0.06$0.03
caching saves on the session with DeepSeek-V4.1-Flash (73%)
$0.59
caching saves on the session with GPT-6 Luna (61%)
$0.17

Why GPT-6 Luna comes out cheaper at list prices

GPT-6 Luna is the cheapest model in OpenAI's GPT-6 family: $0.10 per million input tokens, $0.50 per million output tokens, and $0.01 per million cache hits. DeepSeek-V4.1-Flash lists $0.30 input and $1.20 output at DeepSeek's peak rates. On uncached work that is a 3x gap on the large one-off review, $0.06 against $0.02, and 2.8x on the output-heavy generation, $0.11 against $0.04.

The cache narrows it but does not close it. DeepSeek's cache hit is $0.006 per million, 2% of its input price, and OpenAI's is $0.01, 10% of Luna's. DeepSeek lists no fee for writing the cache, while OpenAI charges 1.25x input for a write from GPT-5.6 on, $0.125 on Luna. Because Luna's input is so low, that premium still leaves its writes at $0.05 for the session's 400K written tokens, against $0.12 on DeepSeek.

The session lands at $0.22 on DeepSeek-V4.1-Flash and $0.11 on GPT-6 Luna, a 2x gap, and 110 sessions a month come to $24.42 against $11.55. Cache writes are the largest line on both, 55% of DeepSeek's session and 45% of Luna's, so the write price matters more here than the headline input rate.

Does DeepSeek's off-peak discount change the answer?

It narrows it a lot. The DeepSeek prices on this page are peak rates, and DeepSeek charges 50% less off-peak, cache hits included. Peak runs from 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays, which is seven hours of a weekday and none of the weekend.

Halving every DeepSeek line brings the example session close to what GPT-6 Luna charges, though Luna stays slightly lower. Uncached work still favors Luna, because DeepSeek's discounted input and output rates remain above Luna's $0.10 and $0.50.

The discount is listed on DeepSeek's own price page, so it applies to DeepSeek's API. OpenRouter sends a request for an open-weight model to one of several providers, and each sets its own price, which can differ from DeepSeek's. The tables here use each maker's own price.

Context, output, and long prompts

GPT-6 Luna has a 1.05M context window and DeepSeek-V4.1-Flash 1M, so both can hold a large repository. Luna carries a long-context rule: once a request passes 272K input tokens, OpenAI bills all of it at 2x for input and cache and 1.5x for output. Every request in the example session stays below 200K, so the Luna figures above use its standard rates.

Output limits differ more. DeepSeek lists 384K output tokens against 128K for Luna. OpenAI sets Luna's reasoning effort at medium by default and lets Codex raise it to max, and more effort means more output tokens. Output is 27% of the session on both models here, so effort settings move both totals.

OpenAI calls GPT-6 Luna "our most efficient model for focused, high-volume tasks," and the Codex docs recommend it for focused, repeatable work. DeepSeek presents V4.1 Flash as the smallest model in its new architecture family, with native visual understanding, and says its KV cache needs a quarter of the memory of the previous generation, cutting cache-hit costs for agents.

EveryToken prices GPT-6 Luna at OpenAI's rates from your Codex or OpenCode history, and prices DeepSeek-V4.1-Flash only when you use it through OpenRouter, from OpenRouter's catalog.

Prompt caching

How each maker bills cached tokens

DeepSeek

DeepSeek's disk cache is on by default for every account, with no code changes. DeepSeek lists no fee for writing the cache.

A cache hit costs $0.006 per million tokens on DeepSeek-V4.1-Flash and $0.044 on DeepSeek-V4-Pro at peak rates, and off-peak hours cost 50% less.

Source: DeepSeek API: Context caching

OpenAI

Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.

From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.

On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.

Source: OpenAI: Prompt caching

Your own numbers

See what DeepSeek-V4.1-Flash and GPT-6 Luna really cost you.

everyaitoken reads your OpenRouter, Codex, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is GPT-6 Luna cheaper than DeepSeek-V4.1-Flash?

At published rates, yes. Input costs $0.10 against $0.30, output $0.50 against $1.20, and the example cached session $0.11 against $0.22. DeepSeek's cache hits are cheaper, $0.006 against $0.01 per million, but not by enough to make up the difference.

When is DeepSeek-V4.1-Flash off-peak?

Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. Every other hour is off-peak and costs 50% less, including cache hits.

Where can I use each model?

GPT-6 Luna runs in Codex, OpenRouter, OpenCode, and GitHub Copilot, but not in Codex cloud. DeepSeek-V4.1-Flash is available through OpenRouter and OpenCode and on DeepSeek's own API as deepseek-flash, and it is not offered in Cursor or Copilot.

Does GPT-6 Luna have a long-context surcharge?

Yes. A Luna request that passes 272K input tokens is billed in full at 2x for input and cache and 1.5x for output, not just the part above the line. Below it, the $0.10 input and $0.50 output rates apply.

  • DeepSeek-V4.1-Flash vs GPT-5.6 Luna

    DeepSeek-V4.1-Flash and GPT-5.6 Luna both price a cached coding session at $0.22, from opposite ends: Luna's cheaper input, DeepSeek's cheaper cache hits.

  • DeepSeek-V4.1-Flash vs DeepSeek-V4-Pro

    DeepSeek says DeepSeek-V4.1-Flash outperforms its own DeepSeek-V4-Pro, which costs 4.3x as much per coding session. What Pro still offers, and what comes next.

  • GPT-6 Astra vs GPT-6 Luna

    GPT-6 Astra and GPT-6 Luna sit at opposite ends of OpenAI's lineup, 100x apart per token. What each is for, and where GPT-6 Sol fits between them.

  • GPT-6 Sol vs GPT-6 Luna

    GPT-6 Luna costs a twentieth of GPT-6 Sol per token. What OpenAI and the Codex docs say each tier is for, and what the gap means for coding sessions.

  • GPT-6 Luna vs GPT-5.6 Luna

    GPT-6 Luna halves the input price of GPT-5.6 Luna, which already had an 80% cut, and trims output further. What the move saves, and what Codex suggests.

  • Claude Haiku 4.5 vs GPT-6 Luna

    GPT-6 Luna lists a tenth of Claude Haiku 4.5's rates and a 1.05M context window against 200K. What each costs for sub-agents and cached coding sessions.