Skip to content

Model comparison

Grok Build 0.1 vs GPT-6 Luna: $1.00 or $0.11 a session

GPT-6 Luna charges a tenth of Grok Build 0.1's input price, and a cached coding session costs $0.11 against $1.00. What each model is for, and where it runs.

· Prices as of September 28, 2026

  • Grok Build 0.1

    xAI · Released May 2026 · Preview

    xAI's dedicated agentic coding model, which also answers to the older grok-code-fast ids.

    Grok Build 0.1 facts and comparisons
  • GPT-6 Luna

    OpenAI · Released September 22, 2026

    The cheapest GPT-6 model, pitched for focused, high-volume, repeatable work, including narrower coding tasks.

    GPT-6 Luna facts and comparisons

The short answer

GPT-6 Luna is far cheaper: the example agentic coding session costs $0.11 on it against $1.00 on Grok Build 0.1, and a month of 110 sessions costs $11.55 against $110.00. The gap is narrowest on output-heavy work, 4.8x, because Grok Build 0.1's $2 output rate is 4x Luna's $0.50, while its cache hits cost 20x as much. OpenAI pitches Luna for focused, high-volume, repeatable work, and xAI describes Grok Build 0.1 as a coding model trained for agentic workflows, still in early access with a 256K window.

Choose Grok Build 0.1 if

  • You want a model xAI trained specifically for agentic coding workflows, rather than one pitched for focused, repeatable tasks.
  • You already call the older grok-code-fast ids, which now answer to Grok Build 0.1.
  • You work through OpenRouter or OpenCode and keep prompts under 200K tokens, below Grok Build 0.1's price step.

Choose GPT-6 Luna if

  • Cost comes first: Luna's rates are $0.10 input, $0.01 cached, and $0.50 output per million.
  • You run high-volume or repeatable coding tasks, which OpenAI and the Codex docs name as Luna's focus.
  • You need more than 256K tokens of context, since Luna offers 1.05M.
  • You work in Codex, or in GitHub Copilot, which offers Luna and does not list Grok Build 0.1.

Side by side

Specs and prices

FactGrok Build 0.1GPT-6 Luna
MakerxAIOpenAI
API model idgrok-build-0.1gpt-6-luna
ReleasedMay 2026September 22, 2026
StatusPreviewCurrent
Context window256K tokens1.05M tokens
Max outputNot published128K tokens
Open weightsNoNo
Input, per 1M tokens$1$0.10
Cache hit, per 1M$0.20$0.01
Cache write, per 1M$1 (same as input)$0.125
Output, per 1M tokens$2$0.50
Runs inOpenCode and OpenRouterCodex, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Grok Build 0.1: Once a prompt reaches 200K tokens, every token in the request costs $2 input, $0.40 cached, and $4 output per million. GPT-6 Luna: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadGrok Build 0.1GPT-6 Luna
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$1.00$0.11
Large one-off review, 150K input with no cache hits, 10K output$0.17$0.02
Output-heavy generation, 30K input, 80K output$0.19$0.04
A month of sessions, 110 sessions: 5 a day, 22 working days$110.00$11.55
Where the session’s cost goes
Cache writes$0.40$0.05
Cache reads$0.40$0.02
Uncached input$0.10$0.01
Output$0.10$0.03
caching saves on the session with Grok Build 0.1 (62%)
$1.60
caching saves on the session with GPT-6 Luna (61%)
$0.17

A 9.1x gap on the agentic coding session

GPT-6 Luna charges $0.10 per million input tokens, $0.01 per cache hit, $0.125 per cache write, and $0.50 per million output tokens. Grok Build 0.1 charges $1, $0.20, $1, and $2. Input is 10x dearer on Grok Build 0.1, cache hits 20x, and output 4x.

The example session is mostly cache reads and writes, where the gaps are largest. Reads cost $0.40 on Grok Build 0.1 and $0.02 on Luna, and writes cost $0.40 and $0.05. With input and output added, the session comes to $1.00 against $0.11, a 9.1x gap. Over 110 sessions a month that is $110.00 against $11.55, or $98.45 apart.

Output narrows it. The output-heavy generation costs $0.19 on Grok Build 0.1 and $0.04 on Luna, a 4.8x gap, because output is the one rate where Grok Build 0.1 is within 4x of Luna. The large one-off review falls between the two at 8.5x, $0.17 against $0.02.

Cache writes: a premium on Luna, none on Grok Build 0.1

From GPT-5.6 on, OpenAI bills a cache write at 1.25x input, so a Luna write costs $0.125 per million. xAI lists no write fee, so Grok Build 0.1 bills written tokens at its $1 input rate. The premium barely registers at Luna's prices: writes are 45% of Luna's session and still come to $0.05.

The mechanics differ as well. OpenAI caching is on by default, starts at 1,024 input tokens, allows up to four explicit breakpoints, and keeps a cached prefix reusable for at least 30 minutes after its last use. xAI caches repeated prefixes automatically and says sending the same conversation id with each request raises the hit rate. Caching saves 62% of Grok Build 0.1's uncached session cost and 61% of Luna's.

Context windows and long-prompt pricing

Luna has a 1.05M context window and writes up to 128K tokens of output. Grok Build 0.1 accepts 256K, and xAI publishes no output limit. A repository that needs more than 256K tokens in one request fits on Luna alone.

Both raise prices for long prompts. Once a Grok Build 0.1 prompt reaches 200K tokens, every token in the request costs $2 input, $0.40 cached, and $4 output per million. Luna's requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request. Grok Build 0.1's step comes well inside its own window, and Luna's comes only after Grok Build 0.1's window has already ended.

What xAI and OpenAI built each model for

OpenAI calls Luna "our most efficient model for focused, high-volume tasks," and the Codex docs recommend it for focused, repeatable tasks. Its reasoning effort defaults to medium, and in Codex it goes up to max. It is not available in Codex cloud, and Free and Go plans get it in the Codex app. Luna runs in Codex, OpenCode, OpenRouter, and GitHub Copilot.

xAI describes Grok Build 0.1 as "xAI's coding model, trained specifically for agentic coding workflows," with function calling, structured outputs, and reasoning. xAI announced it as early access in May 2026, and it answers to the older grok-code-fast ids. Among the tools compared here, you reach it through OpenRouter and OpenCode.

EveryToken prices Luna at OpenAI's rates from Codex or OpenCode history, and prices Grok Build 0.1 when it runs through OpenRouter, from OpenRouter's catalog, whereas the tables use xAI's own API price. At these rates the dollar amounts are small, so the share of each session that comes from cache reads is often the more telling number.

Prompt caching

How each maker bills cached tokens

xAI

The xAI API caches repeated prompt prefixes automatically. Sending the same conversation id with each request raises the hit rate.

xAI lists no fee for writing the cache. A cache hit costs $0.50 per million tokens on Grok 4.7 and $0.20 on Grok Build 0.1.

Source: xAI docs: Prompt caching

OpenAI

Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.

From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.

On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.

Source: OpenAI: Prompt caching

Your own numbers

See what Grok Build 0.1 and GPT-6 Luna really cost you.

everyaitoken reads your OpenRouter, Codex, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

How much cheaper is GPT-6 Luna than Grok Build 0.1?

On the example agentic session, $0.11 against $1.00, a 9.1x gap. The gap shrinks to 4.8x on output-heavy work, where Grok Build 0.1's $2 output rate is 4x Luna's $0.50.

What does OpenAI recommend GPT-6 Luna for?

OpenAI calls it its most efficient model for focused, high-volume tasks, and the Codex docs recommend it for focused, repeatable tasks. xAI pitches Grok Build 0.1 at agentic coding workflows instead.

Do both models charge more for long prompts?

Yes, at different lines. Grok Build 0.1 doubles every rate once a prompt reaches 200K tokens. GPT-6 Luna bills requests over 272K input tokens at 2x for input and cache and 1.5x for output.

Where can I use each model?

GPT-6 Luna runs in Codex, OpenCode, OpenRouter, and GitHub Copilot, though not in Codex cloud. Grok Build 0.1 runs through OpenRouter and OpenCode.

  • GPT-6 Astra vs GPT-6 Luna

    GPT-6 Astra and GPT-6 Luna sit at opposite ends of OpenAI's lineup, 100x apart per token. What each is for, and where GPT-6 Sol fits between them.

  • GPT-6 Sol vs GPT-6 Luna

    GPT-6 Luna costs a twentieth of GPT-6 Sol per token. What OpenAI and the Codex docs say each tier is for, and what the gap means for coding sessions.

  • Grok 4.7 vs Grok Build 0.1

    Grok Build 0.1 costs $1.00 per cached coding session against $2.30 on Grok 4.7. Both follow xAI's caching and 200K rules; they differ on window and status.

  • GPT-6 Luna vs GPT-5.6 Luna

    GPT-6 Luna halves the input price of GPT-5.6 Luna, which already had an 80% cut, and trims output further. What the move saves, and what Codex suggests.

  • Claude Haiku 4.5 vs GPT-6 Luna

    GPT-6 Luna lists a tenth of Claude Haiku 4.5's rates and a 1.05M context window against 200K. What each costs for sub-agents and cached coding sessions.

  • DeepSeek-V4.1-Flash vs GPT-6 Luna

    GPT-6 Luna undercuts DeepSeek-V4.1-Flash at list prices, $0.11 against $0.22 for a cached coding session. How cache pricing and off-peak hours move the gap.