Skip to content

OpenAI

GPT-6 Luna: price, context window, and caching

The cheapest GPT-6 model, pitched for focused, high-volume, repeatable work, including narrower coding tasks.

Released September 22, 2026 · Prices as of September 28, 2026

In OpenAI’s words

“Our most efficient model for focused, high-volume tasks.”

OpenAI docs: GPT-6 Luna

What OpenAI says it’s good at

  • Focused, repeatable tasks, as recommended in the Codex docs Source
  • At higher effort, matching GPT-5.6 Sol on OpenAI's factuality evaluation at about a hundredth of its cost Source

Facts

Specs and prices

FactGPT-6 Luna
MakerOpenAI
API model idgpt-6-luna
ReleasedSeptember 22, 2026
StatusCurrent
Context window1.05M tokens
Max output128K tokens
Open weightsNo
Input, per 1M tokens$0.10
Cache hit, per 1M$0.01
Cache write, per 1M$0.125
Output, per 1M tokens$0.50
Runs inCodex, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by the maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. GPT-6 Luna: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.

Good to know

  • The default reasoning effort is medium. In Codex it supports effort up to max.
  • Not available in Codex cloud. Free and Go plans get it in the Codex app.

Cost

What typical work costs

Example token counts at GPT-6 Luna’s published rates. On the agentic session, caching saves $0.17 against billing every token as ordinary input.

Example workload costs for GPT-6 Luna
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$0.11
Large one-off review, 150K input with no cache hits, 10K output$0.02
Output-heavy generation, 30K input, 80K output$0.04
A month of sessions, 110 sessions: 5 a day, 22 working days$11.55

Prompt caching

How OpenAI bills cached tokens

OpenAI

Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.

From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.

On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.

Source: OpenAI: Prompt caching

GPT-6 Luna compared

  • Claude Haiku 4.5 vs GPT-6 Luna

    GPT-6 Luna lists a tenth of Claude Haiku 4.5's rates and a 1.05M context window against 200K. What each costs for sub-agents and cached coding sessions.

  • DeepSeek-V4.1-Flash vs GPT-6 Luna

    GPT-6 Luna undercuts DeepSeek-V4.1-Flash at list prices, $0.11 against $0.22 for a cached coding session. How cache pricing and off-peak hours move the gap.

  • GLM-5.3-Flash vs GPT-6 Luna

    GPT-6 Luna and GLM-5.3-Flash share a $0.50 output rate, yet a cached coding session costs $0.11 on Luna and $0.16 on GLM-5.3-Flash. Cache hits explain why.

  • GPT-6 Astra vs GPT-6 Luna

    GPT-6 Astra and GPT-6 Luna sit at opposite ends of OpenAI's lineup, 100x apart per token. What each is for, and where GPT-6 Sol fits between them.

  • GPT-6 Luna vs Gemini 3.8 Flash

    Gemini 3.8 Flash costs 7.5x as much as GPT-6 Luna per token, and its introductory rates end December 31, 2026. A month of sessions: $11.55 vs $78.38.

  • GPT-6 Luna vs Gemini 3.5 Flash-Lite

    GPT-6 Luna lists a third of Gemini 3.5 Flash-Lite's input rate and a fifth of its output rate. A month of cached coding sessions: $11.55 against $36.85.

  • GPT-6 Sol vs GPT-6 Luna

    GPT-6 Luna costs a twentieth of GPT-6 Sol per token. What OpenAI and the Codex docs say each tier is for, and what the gap means for coding sessions.

  • Grok Build 0.1 vs GPT-6 Luna

    GPT-6 Luna charges a tenth of Grok Build 0.1's input price, and a cached coding session costs $0.11 against $1.00. What each model is for, and where it runs.

  • MiniMax M3 vs GPT-6 Luna

    GPT-6 Luna costs $0.11 per cached coding session against $0.33 on MiniMax M3. Where the 3x gap comes from, and what MiniMax M3 offers in return.

  • GPT-6 Luna vs GPT-5.6 Luna

    GPT-6 Luna halves the input price of GPT-5.6 Luna, which already had an 80% cut, and trims output further. What the move saves, and what Codex suggests.

Your own numbers

See what GPT-6 Luna really costs you.

everyaitoken reads your Codex, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math