Skip to content

Model comparison

DeepSeek-V4.1-Flash and GPT-5.6 Luna tie on session cost

DeepSeek-V4.1-Flash and GPT-5.6 Luna both price a cached coding session at $0.22, from opposite ends: Luna's cheaper input, DeepSeek's cheaper cache hits.

· Prices as of September 28, 2026

  • DeepSeek-V4.1-Flash

    DeepSeek · Released September 10, 2026

    DeepSeek's small, low-cost, natively multimodal model on its new architecture, which DeepSeek says outperforms its own DeepSeek-V4-Pro.

    DeepSeek-V4.1-Flash facts and comparisons
  • GPT-5.6 Luna

    OpenAI · Released July 9, 2026 · Previous generation

    The low-cost GPT-5.6 tier for cost-sensitive, high-volume work, which roughly corresponds to earlier GPT-5 nano models.

    GPT-5.6 Luna facts and comparisons

The short answer

DeepSeek-V4.1-Flash and GPT-5.6 Luna cost the same $0.22 for the example agentic coding session, and 110 sessions a month differ by $0.22, because Luna's lower input price and DeepSeek's lower cache-hit price cancel out. GPT-5.6 Luna is a previous-generation model that Codex suggests replacing with GPT-6 Luna, so it suits teams already running it in Codex, Cursor, or GitHub Copilot. DeepSeek-V4.1-Flash is the current model of the two, with open weights and a 50% off-peak discount that tips the session its way.

Choose DeepSeek-V4.1-Flash if

  • You can schedule work outside DeepSeek's peak hours, when it charges 50% less.
  • Your agent resends a large cached context: a DeepSeek hit costs $0.006 per million against $0.02, under a third of the price.
  • You want a current model with open weights under the MIT license rather than a previous-generation tier.
  • You need responses longer than 128K tokens, since DeepSeek lists 384K.

Choose GPT-5.6 Luna if

  • Your work is mostly uncached: Luna's $0.20 input makes the large one-off review $0.04 against $0.06.
  • You already use it in Codex, Cursor, or GitHub Copilot and want to keep a working setup while you evaluate GPT-6 Luna.
  • You want explicit cache breakpoints, up to four, which OpenAI supports from GPT-5.6 on.
  • Your team works during DeepSeek's peak hours, when its full rates apply.

Side by side

Specs and prices

FactDeepSeek-V4.1-FlashGPT-5.6 Luna
MakerDeepSeekOpenAI
API model iddeepseek-flashgpt-5.6-luna
ReleasedSeptember 10, 2026July 9, 2026
StatusCurrentPrevious generation
Context window1M tokens1.05M tokens
Max output384K tokens128K tokens
Open weightsYesNo
Input, per 1M tokens$0.30$0.20
Cache hit, per 1M$0.006$0.02
Cache write, per 1M$0.30 (same as input)$0.25
Output, per 1M tokens$1.20$1.20
Runs inOpenCode and OpenRouterCodex, Cursor, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. DeepSeek-V4.1-Flash: Prices are DeepSeek's peak rates. Off-peak hours cost 50% less: peak is 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. GPT-5.6 Luna: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadDeepSeek-V4.1-FlashGPT-5.6 Luna
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$0.22$0.22
Large one-off review, 150K input with no cache hits, 10K output$0.06$0.04
Output-heavy generation, 30K input, 80K output$0.11$0.10
A month of sessions, 110 sessions: 5 a day, 22 working days$24.42$24.20
Where the session’s cost goes
Cache writes$0.12$0.10
Cache reads$0.01$0.04
Uncached input$0.03$0.02
Output$0.06$0.06
caching saves on the session with DeepSeek-V4.1-Flash (73%)
$0.59
caching saves on the session with GPT-5.6 Luna (61%)
$0.34

How two different price lists reach the same $0.22

The rate cards look nothing alike. GPT-5.6 Luna charges $0.20 per million input tokens, $0.25 per million for a cache write, and $0.02 per million for a cache hit. DeepSeek-V4.1-Flash charges $0.30 for input, bills cache writes as ordinary input because DeepSeek lists no fee for writing its cache, and charges $0.006 for a hit. Output is $1.20 per million on both.

Line by line, the example session trades evenly. Luna saves $0.02 on cache writes, which are 45% of its session, and $0.01 on fresh input. DeepSeek saves $0.03 on cache reads, where 2M tokens cost $0.01 against $0.04. Output costs $0.06 on each. Both sessions round to $0.22, and 110 of them come to $24.42 on DeepSeek and $24.20 on Luna.

Uncached work tilts toward Luna. The large one-off review costs $0.04 against $0.06, a 1.5x gap that is just the input ratio, and the output-heavy generation costs $0.10 against $0.11. The more of its context an agent reads from the cache, the better DeepSeek's side looks; the more fresh input it sends, the better Luna's does.

GPT-5.6 Luna is now a previous-generation model

OpenAI released GPT-5.6 Luna on July 9, 2026 as the low-cost GPT-5.6 tier, which roughly corresponds to earlier GPT-5 nano models, and cut its price by 80% on July 30, 2026. It is now a previous generation. Codex suggests moving to GPT-6 Luna, and GPT-5.6 Luna stays available in the API.

That matters for a long-term choice. OpenAI describes GPT-5.6 Luna as a "GPT-5.6 model optimized for cost-sensitive workloads" and says the GPT-5.6 family reaches flagship-level performance with fewer output tokens. Its successor, GPT-6 Luna, lists lower rates, and Codex points new work there.

DeepSeek-V4.1-Flash is the newer model, released on September 10, 2026. DeepSeek introduces it as "the smallest model in our new architecture family, with native visual understanding" and says tests by several parties put it ahead of DeepSeek-V4-Pro on performance, cost, speed, and total runtime.

Off-peak hours, OpenRouter, and open weights

The DeepSeek rates in these tables are peak rates, which apply from 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, excluding Chinese public holidays. At all other times DeepSeek charges 50% less. With the session tied at peak, a session run off-peak costs about half as much on DeepSeek as on GPT-5.6 Luna.

OpenRouter lists both models. For an open-weight model like DeepSeek-V4.1-Flash, it sends the request to one of several providers, whose prices can differ from DeepSeek's own; the tables use each maker's own price. GPT-5.6 Luna is also in Codex, Cursor, OpenCode, and GitHub Copilot, while DeepSeek-V4.1-Flash is in OpenCode but not in Cursor or Copilot.

DeepSeek publishes the model's weights under the MIT license, so you can also run it yourself. This page does not estimate hosting costs, which depend on hardware and usage. EveryToken prices GPT-5.6 Luna from your Codex, Cursor, or OpenCode history at OpenAI's rates, and prices DeepSeek-V4.1-Flash when you use it through OpenRouter.

Prompt caching

How each maker bills cached tokens

DeepSeek

DeepSeek's disk cache is on by default for every account, with no code changes. DeepSeek lists no fee for writing the cache.

A cache hit costs $0.006 per million tokens on DeepSeek-V4.1-Flash and $0.044 on DeepSeek-V4-Pro at peak rates, and off-peak hours cost 50% less.

Source: DeepSeek API: Context caching

OpenAI

Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.

From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.

On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.

Source: OpenAI: Prompt caching

Your own numbers

See what DeepSeek-V4.1-Flash and GPT-5.6 Luna really cost you.

everyaitoken reads your OpenRouter, Codex, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Do DeepSeek-V4.1-Flash and GPT-5.6 Luna really cost the same?

On the example cached session, yes: both come to $0.22 at published rates, DeepSeek's at its peak. They differ on other work: a large uncached review costs $0.06 on DeepSeek and $0.04 on Luna. Off-peak, DeepSeek charges 50% less.

Should I move from GPT-5.6 Luna to GPT-6 Luna?

Codex suggests that move, and GPT-5.6 Luna stays available in the API meanwhile. The GPT-6 Luna model page lists its rates and limits for comparison.

Which model has cheaper cache hits?

DeepSeek-V4.1-Flash: $0.006 per million, 2% of its input price, against $0.02 on GPT-5.6 Luna, 10% of input. OpenAI also charges 1.25x input to write the cache from GPT-5.6 on, while DeepSeek lists no write fee.

What is the output limit on each?

DeepSeek-V4.1-Flash lists 384K output tokens and GPT-5.6 Luna 128K. Their context windows are 1M and 1.05M tokens.

  • DeepSeek-V4.1-Flash vs DeepSeek-V4-Pro

    DeepSeek says DeepSeek-V4.1-Flash outperforms its own DeepSeek-V4-Pro, which costs 4.3x as much per coding session. What Pro still offers, and what comes next.

  • DeepSeek-V4.1-Flash vs GPT-6 Luna

    GPT-6 Luna undercuts DeepSeek-V4.1-Flash at list prices, $0.11 against $0.22 for a cached coding session. How cache pricing and off-peak hours move the gap.

  • GPT-5.6 Terra vs GPT-5.6 Luna

    GPT-5.6 Terra costs 10x GPT-5.6 Luna on every rate after OpenAI's July 30 price cuts. What each tier is for, and why Codex now sends them different ways.

  • GPT-6 Luna vs GPT-5.6 Luna

    GPT-6 Luna halves the input price of GPT-5.6 Luna, which already had an 80% cut, and trims output further. What the move saves, and what Codex suggests.

  • Claude Haiku 4.5 vs GPT-5.6 Luna

    After an 80% price cut, GPT-5.6 Luna costs a fifth of Claude Haiku 4.5 for input. A month of cached coding sessions: $24.20 against $132.00.

  • DeepSeek-V4.1-Flash vs Gemini 3.8 Flash

    DeepSeek-V4.1-Flash costs $0.22 for a cached coding session that costs $0.71 on Gemini 3.8 Flash, and Gemini's introductory rates end December 31, 2026.