Model comparison
Upgrading from GPT-5.6 Luna to GPT-6 Luna: what changes
GPT-6 Luna halves the input price of GPT-5.6 Luna, which already had an 80% cut, and trims output further. What the move saves, and what Codex suggests.
· Prices as of September 28, 2026
GPT-6 Luna
OpenAI · Released September 22, 2026
The cheapest GPT-6 model, pitched for focused, high-volume, repeatable work, including narrower coding tasks.
GPT-6 Luna facts and comparisonsGPT-5.6 Luna
OpenAI · Released July 9, 2026 · Previous generation
The low-cost GPT-5.6 tier for cost-sensitive, high-volume work, which roughly corresponds to earlier GPT-5 nano models.
GPT-5.6 Luna facts and comparisons
The short answer
GPT-6 Luna costs half as much as GPT-5.6 Luna for input and caching and 58% less for output, so the example agentic coding session drops from $0.22 to $0.11. Codex suggests the move, while GPT-5.6 Luna stays available in the API. At these prices the gap is $12.65 over 110 sessions, so it matters most in high-volume pipelines.
Choose GPT-6 Luna if
- You run high-volume jobs where halving the input rate, from $0.20 to $0.10 per million tokens, adds up.
- Your work is output-heavy: output costs $0.50 per million on GPT-6 Luna against $1.20 on GPT-5.6 Luna.
- You use Codex, which suggests moving to GPT-6 Luna, or the Codex app on a ChatGPT Free or Go plan.
- You want reasoning effort up to max, which GPT-6 Luna supports in Codex.
Choose GPT-5.6 Luna if
- You have a tuned pipeline on GPT-5.6 Luna and would rather not retest it yet, since it stays available in the API.
- You pick models in Cursor, whose model list includes GPT-5.6 Luna, and you have not confirmed GPT-6 Luna there.
Side by side
Specs and prices
| Fact | GPT-6 Luna | GPT-5.6 Luna |
|---|---|---|
| Maker | OpenAI | OpenAI |
| API model id | gpt-6-luna | gpt-5.6-luna |
| Released | September 22, 2026 | July 9, 2026 |
| Status | Current | Previous generation |
| Context window | 1.05M tokens | 1.05M tokens |
| Max output | 128K tokens | 128K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $0.10 | $0.20 |
| Cache hit, per 1M | $0.01 | $0.02 |
| Cache write, per 1M | $0.125 | $0.25 |
| Output, per 1M tokens | $0.50 | $1.20 |
| Runs in | Codex, OpenCode, OpenRouter, and GitHub Copilot | Codex, Cursor, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. GPT-6 Luna: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request. GPT-5.6 Luna: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | GPT-6 Luna | GPT-5.6 Luna |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.11 | $0.22 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.02 | $0.04 |
| Output-heavy generation, 30K input, 80K output | $0.04 | $0.10 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $11.55 | $24.20 |
| Where the session’s cost goes | ||
| Cache writes | $0.05 | $0.10 |
| Cache reads | $0.02 | $0.04 |
| Uncached input | $0.01 | $0.02 |
| Output | $0.03 | $0.06 |
- caching saves on the session with GPT-6 Luna (61%)
- $0.17
- caching saves on the session with GPT-5.6 Luna (61%)
- $0.34
How much cheaper is GPT-6 Luna than GPT-5.6 Luna?
GPT-5.6 Luna charges $0.20 per million input tokens, $0.02 per million cached tokens, $0.25 per million cache-write tokens, and $1.20 per million output tokens. GPT-6 Luna charges $0.10, $0.01, $0.125, and $0.50. Input and caching are exactly 2x apart, and output is 2.4x apart.
That uneven output ratio shows up in the workloads. The session and the uncached review both come out at 2x: $0.22 against $0.11, and $0.04 against $0.02. The output-heavy generation shows 2.5x, $0.10 against $0.04, though at these amounts rounding to the cent moves the ratio. The per-token output rates, 2.4x apart, are the steadier guide.
A month of 110 sessions costs $24.20 on GPT-5.6 Luna and $11.55 on GPT-6 Luna, a difference of $12.65 at published API rates. For one developer that is small. For a service that sends thousands of Luna requests a day, the 2x input gap and the 2.4x output gap are what scale.
GPT-5.6 Luna already had an 80% price cut
GPT-5.6 Luna launched on July 9, 2026, and OpenAI cut its price by 80% on July 30, 2026. The rates above are the post-cut rates. GPT-6 Luna, released on September 22, 2026, still comes in at half of them for input and less than half for output.
The positioning moved with the new generation. OpenAI described GPT-5.6 Luna as optimized for cost-sensitive workloads, roughly where earlier GPT-5 nano models sat. It calls GPT-6 Luna its most efficient model for focused, high-volume tasks and pitches it for repeatable work, including narrower coding tasks. Codex's own docs list it for focused, repeatable tasks.
OpenAI also claims that at higher effort, GPT-6 Luna matches GPT-5.6 Sol on its factuality evaluation at about a hundredth of the cost. That claim compares Luna with the previous flagship, not with GPT-5.6 Luna, and it measures factual accuracy rather than coding.
Codex, caching, and context for the two Luna models
Codex suggests moving from GPT-5.6 Luna to GPT-6 Luna, and GPT-5.6 Luna stays available in the API. GPT-6 Luna defaults to medium reasoning effort and supports effort up to max in Codex. It is not available in Codex cloud, and ChatGPT Free and Go plans get it in the Codex app.
Caching works identically, because both come from GPT-5.6 or later: a write costs 1.25x input, a hit 0.1x, and caching starts at 1,024 input tokens. On the example session caching saves $0.34 on GPT-5.6 Luna and $0.17 on GPT-6 Luna, 61% on each. Context and output limits match at 1.05M and 128K, and both switch to higher rates above 272K input tokens.
At fractions of a cent per request, it is easy to lose track of where tokens go. EveryToken reads your local Codex and OpenCode history and prices each request at OpenAI's API rates, so a switch from one Luna to the other shows up in real numbers.
Prompt caching
How OpenAI bills cached tokens
OpenAI
Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.
From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.
On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.
Source: OpenAI: Prompt caching
Your own numbers
See what GPT-6 Luna and GPT-5.6 Luna really cost you.
everyaitoken reads your Codex, OpenCode, OpenRouter, and Cursor history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is GPT-6 Luna half the price of GPT-5.6 Luna?
For input, cache writes, and cache hits, yes. Output drops further, from $1.20 to $0.50 per million tokens, which is 58% less.
Is GPT-5.6 Luna being retired?
Not from the API, where it stays available. Codex suggests moving to GPT-6 Luna instead.
What role does GPT-6 Luna take in OpenAI's lineup?
It is the low-cost tier of GPT-6, the role GPT-5.6 Luna held in the GPT-5.6 family. OpenAI placed GPT-5.6 Luna roughly where earlier GPT-5 nano models sat.
Does the 272K long-context rule apply to both Luna models?
Yes. On both, requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.
Sources
- OpenAI: API pricing
- OpenAI docs: GPT-6 Luna
- OpenAI: Introducing GPT-6 Sol and Luna
- OpenAI: API changelog
- Codex docs: Models
- OpenRouter: GPT-6 Luna
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- OpenAI docs: GPT-5.6 Luna
- OpenAI: Using GPT-5.6
- Cursor docs: Models and pricing
- OpenRouter: GPT-5.6 Luna
- OpenAI: Prompt caching