Skip to content

Model comparison

GPT-6 Luna or Gemini 3.8 Flash for high-volume coding?

Gemini 3.8 Flash costs 7.5x as much as GPT-6 Luna per token, and its introductory rates end December 31, 2026. A month of sessions: $11.55 vs $78.38.

· Prices as of September 28, 2026

  • GPT-6 Luna

    OpenAI · Released September 22, 2026

    The cheapest GPT-6 model, pitched for focused, high-volume, repeatable work, including narrower coding tasks.

    GPT-6 Luna facts and comparisons
  • Gemini 3.8 Flash

    Google · Released September 2, 2026

    Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.

    Gemini 3.8 Flash facts and comparisons

The short answer

GPT-6 Luna costs far less: its rates are 7.5x below Gemini 3.8 Flash's, and a month of example agentic coding sessions comes to $11.55 against $78.38. The two sit in different tiers, since OpenAI pitches Luna for focused, high-volume work while Google pitches Gemini 3.8 Flash for long-horizon software engineering and autonomous agents. Gemini 3.8 Flash is on introductory rates through December 31, 2026, after which its listed prices double.

Choose GPT-6 Luna if

  • Cost comes first: $0.10 input and $0.50 output per million, against $0.75 and $3.75.
  • Your tasks are narrow and repeatable, the kind the Codex docs recommend Luna for.
  • You need long responses: Luna writes up to 128K tokens against Gemini 3.8 Flash's 65.5K.
  • Your team uses the Codex app on Free or Go plans, which include Luna.

Choose Gemini 3.8 Flash if

  • Your work is long-horizon software engineering or multi-file refactoring, which Google names as Gemini 3.8 Flash's focus.
  • You run Gemini CLI with a Gemini API key or Vertex AI, and its default auto model already uses Gemini 3.8 Flash as the Flash half.
  • You want to start free: the Gemini API free tier covers its input, output, and caching.
  • Your agents read untrusted content, and Google's claim of better robustness against prompt injection matters to you.

Side by side

Specs and prices

FactGPT-6 LunaGemini 3.8 Flash
MakerOpenAIGoogle
API model idgpt-6-lunagemini-3.8-flash
ReleasedSeptember 22, 2026September 2, 2026
StatusCurrentCurrent
Context window1.05M tokens1.05M tokens
Max output128K tokens65.5K tokens
Open weightsNoNo
Input, per 1M tokens$0.10$0.75
Cache hit, per 1M$0.01$0.075
Cache write, per 1M$0.125$0.75 (same as input)
Output, per 1M tokens$0.50$3.75
Runs inCodex, OpenCode, OpenRouter, and GitHub CopilotCursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. GPT-6 Luna: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadGPT-6 LunaGemini 3.8 Flash
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$0.11$0.71
Large one-off review, 150K input with no cache hits, 10K output$0.02$0.15
Output-heavy generation, 30K input, 80K output$0.04$0.32
A month of sessions, 110 sessions: 5 a day, 22 working days$11.55$78.38
Where the session’s cost goes
Cache writes$0.05$0.30
Cache reads$0.02$0.15
Uncached input$0.01$0.08
Output$0.03$0.19
caching saves on the session with GPT-6 Luna (61%)
$0.17
caching saves on the session with Gemini 3.8 Flash (66%)
$1.35

Two different tiers, a 6.8x monthly gap

GPT-6 Luna is the low-cost end of the GPT-6 family. OpenAI calls it "our most efficient model for focused, high-volume tasks." Gemini 3.8 Flash is Google's newest Flash, described as "our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows." Comparing them sets OpenAI's budget tier against Google's main Flash tier.

The price difference follows. Luna charges $0.10 per million input tokens, $0.01 per million cache hits, and $0.50 per million output tokens. Gemini 3.8 Flash charges $0.75, $0.075, and $3.75, each 7.5x higher. The large one-off review costs $0.02 against $0.15.

Over 110 example sessions a month, Luna costs $11.55 and Gemini 3.8 Flash $78.38, a 6.8x gap and $66.83 apart. The monthly ratio is lower than the per-token ratio because cache writes are closer: OpenAI charges 1.25x input for a write, $0.125 per million on Luna, while Google bills written tokens as ordinary input at $0.75. That write gap is 6x rather than 7.5x.

What January 1, 2027 does to the gap

Gemini 3.8 Flash's current rates are introductory and run through December 31, 2026. From January 1, 2027, Google lists $1.50 per million input tokens, $0.15 per million cached tokens, and $7.50 per million output tokens, twice today's prices. Luna's price notes list no promotion, so on current information the gap between the two roughly doubles in the new year.

The session figures in this post use today's rates. If you are sizing a budget that runs into 2027, redo the Gemini side at the higher rates before comparing.

Thinking defaults, output limits, and caching

Both default to a medium setting: Luna's reasoning effort and Gemini 3.8 Flash's thinking level. Gemini CLI sends high for Gemini 3.8 Flash, and Codex lets Luna go up to max effort. Higher settings write more tokens, and output is 26% of Gemini 3.8 Flash's session cost here and 27% of Luna's.

Their context windows match at 1.05M tokens. Luna writes up to 128K tokens per response, twice Gemini 3.8 Flash's 65.5K. Luna's rates rise for any request over 272K input tokens: 2x for input and cache and 1.5x for output, for the whole request.

Caching saves 66% of the uncached session cost on Gemini 3.8 Flash, $1.35, and 61% on Luna, $0.17. Google's implicit caching is automatic but not assured to hit, and its explicit caching adds a storage charge of $0.50 to $1 per million tokens per hour on Flash models. OpenAI's cached prefixes stay reusable for at least 30 minutes after their last use.

Prompt caching

How each maker bills cached tokens

OpenAI

Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.

From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.

On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.

Source: OpenAI: Prompt caching

Google

Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.

Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.

On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.

Source: Google: Context caching

Your own numbers

See what GPT-6 Luna and Gemini 3.8 Flash really cost you.

everyaitoken reads your Codex, OpenCode, OpenRouter, Gemini CLI, and Cursor history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

How much cheaper is GPT-6 Luna than Gemini 3.8 Flash?

7.5x per token on input, output, and cache hits, and 6.8x over a month of example agentic sessions, $11.55 against $78.38. The gap is set to widen when Gemini 3.8 Flash's introductory rates end.

Is there a free way to use Gemini 3.8 Flash?

Google's free tier on the Gemini API includes its input, output, and caching. Paid use costs $0.75 input and $3.75 output per million tokens until December 31, 2026.

Which of the two is aimed at complex coding?

Google positions Gemini 3.8 Flash for long-horizon software engineering, multi-file refactoring, and autonomous agents. OpenAI positions GPT-6 Luna for focused, repeatable, high-volume tasks, and the Codex docs recommend GPT-6 Sol for complex coding.

Where can I see both models' costs together?

EveryToken reads your local Codex and Gemini CLI history, prices each request at API rates, and shows cost and cache savings per model side by side.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash

    Gemini 3.8 Flash costs about half of Gemini 3.5 Flash on introductory rates. What changes on January 1, 2027, and which one Gemini CLI picks for you.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash-Lite

    Google recommends Gemini 3.8 Flash and Gemini 3.5 Flash-Lite together for new projects. Where each fits, the 2.1x session gap, and what changes in 2027.

  • Gemini 3.8 Flash vs Gemini 3.1 Pro Preview

    Gemini CLI's auto model uses both Gemini 3.8 Flash and Gemini 3.1 Pro Preview. What each costs, why Flash rates double in 2027, and the Pro's preview status.

  • Gemini 3.8 Flash vs Gemini 3.7 Flash

    Gemini 3.8 Flash and Gemini 3.7 Flash cost the same, and both rates double on January 1, 2027. What differs is Gemini CLI, Cursor, and Google's own pitch.

  • GPT-5.6 Terra vs Gemini 3.8 Flash

    GPT-5.6 Terra costs 3.1x as much as Gemini 3.8 Flash on a cached coding session, $2.20 against $0.71, and its listed 2027 rates stay below Terra's.

  • GPT-6 Astra vs Gemini 3.8 Flash

    GPT-6 Astra and Gemini 3.8 Flash launched a day apart at opposite ends of the price range, 14.8x apart on a cached coding session. What each is built for.