Model comparison
GPT-6 Astra vs GPT-6 Luna: the 100x spread in one lineup
GPT-6 Astra and GPT-6 Luna sit at opposite ends of OpenAI's lineup, 100x apart per token. What each is for, and where GPT-6 Sol fits between them.
· Prices as of September 28, 2026
GPT-6 Astra
OpenAI · Released September 3, 2026
OpenAI's top GPT-6 model and its most expensive, aimed at the hardest long-running work that spans many tools, including coding.
GPT-6 Astra facts and comparisonsGPT-6 Luna
OpenAI · Released September 22, 2026
The cheapest GPT-6 model, pitched for focused, high-volume, repeatable work, including narrower coding tasks.
GPT-6 Luna facts and comparisons
The short answer
GPT-6 Astra costs 100x as much as GPT-6 Luna per token, so one example agentic coding session comes to $10.50 on Astra and $0.11 on Luna. OpenAI built Astra for the hardest end-to-end work and Luna for focused, high-volume tasks, so they suit different jobs rather than the same one. For complex coding, the Codex docs recommend neither of them but GPT-6 Sol, the middle tier.
Choose GPT-6 Astra if
- You take on the hardest end-to-end work, which OpenAI says GPT-6 Astra is built for.
- Your jobs span many tools over a long run, and you want the model OpenAI says stays coherent during long tasks.
- You leave Codex CLI on its default, and its bundled model list defaults to Astra.
Choose GPT-6 Luna if
- You run focused, repeatable tasks at high volume, the work OpenAI and the Codex docs name for GPT-6 Luna.
- Cost per request shapes your design: Luna charges $0.10 per million input tokens.
- You are on a ChatGPT Free or Go plan, which OpenAI lists as getting GPT-6 Luna in the Codex app.
- You want to push reasoning effort up to max in Codex without large output costs.
Side by side
Specs and prices
| Fact | GPT-6 Astra | GPT-6 Luna |
|---|---|---|
| Maker | OpenAI | OpenAI |
| API model id | gpt-6-astra | gpt-6-luna |
| Released | September 3, 2026 | September 22, 2026 |
| Status | Current | Current |
| Context window | 1.05M tokens | 1.05M tokens |
| Max output | 128K tokens | 128K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $10 | $0.10 |
| Cache hit, per 1M | $1 | $0.01 |
| Cache write, per 1M | $12.50 | $0.125 |
| Output, per 1M tokens | $50 | $0.50 |
| Runs in | Codex, OpenCode, OpenRouter, and GitHub Copilot | Codex, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. GPT-6 Astra: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request. GPT-6 Luna: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | GPT-6 Astra | GPT-6 Luna |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $10.50 | $0.11 |
| Large one-off review, 150K input with no cache hits, 10K output | $2.00 | $0.02 |
| Output-heavy generation, 30K input, 80K output | $4.30 | $0.04 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $1,155.00 | $11.55 |
| Where the session’s cost goes | ||
| Cache writes | $5.00 | $0.05 |
| Cache reads | $2.00 | $0.02 |
| Uncached input | $1.00 | $0.01 |
| Output | $2.50 | $0.03 |
- caching saves on the session with GPT-6 Astra (62%)
- $17.00
- caching saves on the session with GPT-6 Luna (61%)
- $0.17
One GPT-6 Astra session costs about a month of GPT-6 Luna
At the top of the lineup, GPT-6 Astra lists $10 per million input tokens and $50 per million output tokens. At the bottom, GPT-6 Luna lists $0.10 and $0.50. Writing to the cache costs $12.50 per million on Astra and $0.125 on Luna, and reading from it $1 and $0.01. Every rate differs by 100x.
Put in workload terms, one example session on Astra, at $10.50, costs nearly as much as a month of 110 sessions on Luna, at $11.55. Across the month Astra comes to $1,155.00, a difference of $1,143.45 at published API rates.
Rounding blurs the ratio on single workloads. Luna's session rounds to $0.11 and its generation to $0.04, so the ratios read 95.5x and 107.5x rather than 100x. The per-token rates are the reliable comparison, and they sit exactly 100x apart.
What OpenAI built each end of the GPT-6 lineup to do
OpenAI calls Astra its most capable model, built for the hardest end-to-end work, and positions it for long-running work that spans many tools, including coding. It claims Astra stays coherent during long tasks better than GPT-5.6 Sol and earlier models, and follows general instructions more closely than previous models.
Luna sits at the other end. OpenAI's description is its most efficient model for focused, high-volume tasks, and the Codex docs point to it for focused, repeatable jobs. OpenAI also says that at higher effort Luna matches GPT-5.6 Sol on its factuality evaluation at about a hundredth of the cost. None of these claims put Luna and Astra on the same kind of task.
Between them sits GPT-6 Sol, which the Codex docs recommend for complex coding. Weighing Astra against Luna usually comes down to which jobs belong at each end, and whether the rest belong on Sol.
Reasoning effort, Codex defaults, and caching
Both models run in Codex, OpenCode, OpenRouter, and GitHub Copilot, and neither is available in Codex cloud. Astra is the default in Codex CLI's bundled model list, version 0.158.0, and starts there at low reasoning effort. Luna defaults to medium and supports effort up to max in Codex.
Effort matters for cost because reasoning tokens are billed as output. Astra at low effort and Luna at max effort sit at opposite settings, yet output still costs $50 per million on one and $0.50 on the other. Both share a 1.05M context window and 128K of output, and requests past 272K input tokens cost 2x for input and cache and 1.5x for output on both.
Caching follows the same rules on both: writes at 1.25x input, hits at 0.1x, and a prefix that stays reusable for at least 30 minutes after its last use. On the example session caching saves $17.00 on Astra, or 62%, and $0.17 on Luna, or 61%. To see how your own work divides between the tiers, EveryToken prices every local Codex and OpenCode request at OpenAI's API rates.
Prompt caching
How OpenAI bills cached tokens
OpenAI
Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.
From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.
On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.
Source: OpenAI: Prompt caching
Your own numbers
See what GPT-6 Astra and GPT-6 Luna really cost you.
everyaitoken reads your Codex, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Does GPT-6 Astra really cost 100 times as much as GPT-6 Luna?
Per token, yes, on input, output, cache writes, and cache hits. Workload totals round to the cent, so the example session shows $10.50 against $0.11 rather than an exact 100x.
Which GPT-6 model does Codex use by default?
Codex CLI's bundled model list, version 0.158.0, makes GPT-6 Astra the default at low reasoning effort. The Codex docs recommend GPT-6 Sol for complex coding and GPT-6 Luna for focused, repeatable tasks, so the default is worth checking.
Can I use GPT-6 Astra or GPT-6 Luna in Codex cloud?
No. Both run in Codex, and OpenAI lists Astra for the app, the CLI, and the IDE extension, but neither is offered in Codex cloud.
Do GPT-6 Astra and GPT-6 Luna have the same context window?
Yes, 1.05M tokens on each, with output capped at 128K. OpenAI lists Astra as accepting up to 922K input tokens of that window.
Sources
- OpenAI: API pricing
- OpenAI docs: GPT-6 Astra
- OpenAI: Using the latest model
- OpenAI: API changelog
- Codex docs: Models
- Codex CLI: bundled model catalog
- OpenRouter: GPT-6 Astra
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- OpenAI docs: GPT-6 Luna
- OpenAI: Introducing GPT-6 Sol and Luna
- OpenRouter: GPT-6 Luna
- OpenAI: Prompt caching