OpenAI
GPT-6 Luna: price, context window, and caching
The cheapest GPT-6 model, pitched for focused, high-volume, repeatable work, including narrower coding tasks.
Released September 22, 2026 · Prices as of September 28, 2026
In OpenAI’s words
“Our most efficient model for focused, high-volume tasks.”
Facts
Specs and prices
| Fact | GPT-6 Luna |
|---|---|
| Maker | OpenAI |
| API model id | gpt-6-luna |
| Released | September 22, 2026 |
| Status | Current |
| Context window | 1.05M tokens |
| Max output | 128K tokens |
| Open weights | No |
| Input, per 1M tokens | $0.10 |
| Cache hit, per 1M | $0.01 |
| Cache write, per 1M | $0.125 |
| Output, per 1M tokens | $0.50 |
| Runs in | Codex, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by the maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. GPT-6 Luna: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.
Good to know
- The default reasoning effort is medium. In Codex it supports effort up to max.
- Not available in Codex cloud. Free and Go plans get it in the Codex app.
Cost
What typical work costs
Example token counts at GPT-6 Luna’s published rates. On the agentic session, caching saves $0.17 against billing every token as ordinary input.
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.11 |
|---|---|
| Large one-off review, 150K input with no cache hits, 10K output | $0.02 |
| Output-heavy generation, 30K input, 80K output | $0.04 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $11.55 |
Prompt caching
How OpenAI bills cached tokens
OpenAI
Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.
From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.
On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.
Source: OpenAI: Prompt caching
GPT-6 Luna compared
Claude Haiku 4.5 vs GPT-6 Luna
GPT-6 Luna lists a tenth of Claude Haiku 4.5's rates and a 1.05M context window against 200K. What each costs for sub-agents and cached coding sessions.
DeepSeek-V4.1-Flash vs GPT-6 Luna
GPT-6 Luna undercuts DeepSeek-V4.1-Flash at list prices, $0.11 against $0.22 for a cached coding session. How cache pricing and off-peak hours move the gap.
GLM-5.3-Flash vs GPT-6 Luna
GPT-6 Luna and GLM-5.3-Flash share a $0.50 output rate, yet a cached coding session costs $0.11 on Luna and $0.16 on GLM-5.3-Flash. Cache hits explain why.
GPT-6 Astra vs GPT-6 Luna
GPT-6 Astra and GPT-6 Luna sit at opposite ends of OpenAI's lineup, 100x apart per token. What each is for, and where GPT-6 Sol fits between them.
GPT-6 Luna vs Gemini 3.8 Flash
Gemini 3.8 Flash costs 7.5x as much as GPT-6 Luna per token, and its introductory rates end December 31, 2026. A month of sessions: $11.55 vs $78.38.
GPT-6 Luna vs Gemini 3.5 Flash-Lite
GPT-6 Luna lists a third of Gemini 3.5 Flash-Lite's input rate and a fifth of its output rate. A month of cached coding sessions: $11.55 against $36.85.
GPT-6 Sol vs GPT-6 Luna
GPT-6 Luna costs a twentieth of GPT-6 Sol per token. What OpenAI and the Codex docs say each tier is for, and what the gap means for coding sessions.
Grok Build 0.1 vs GPT-6 Luna
GPT-6 Luna charges a tenth of Grok Build 0.1's input price, and a cached coding session costs $0.11 against $1.00. What each model is for, and where it runs.
MiniMax M3 vs GPT-6 Luna
GPT-6 Luna costs $0.11 per cached coding session against $0.33 on MiniMax M3. Where the 3x gap comes from, and what MiniMax M3 offers in return.
GPT-6 Luna vs GPT-5.6 Luna
GPT-6 Luna halves the input price of GPT-5.6 Luna, which already had an 80% cut, and trims output further. What the move saves, and what Codex suggests.
Your own numbers
See what GPT-6 Luna really costs you.
everyaitoken reads your Codex, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.