Model comparison
GPT-6 Luna or Gemini 3.8 Flash for high-volume coding?
Gemini 3.8 Flash costs 7.5x as much as GPT-6 Luna per token, and its introductory rates end December 31, 2026. A month of sessions: $11.55 vs $78.38.
· Prices as of September 28, 2026
GPT-6 Luna
OpenAI · Released September 22, 2026
The cheapest GPT-6 model, pitched for focused, high-volume, repeatable work, including narrower coding tasks.
GPT-6 Luna facts and comparisonsGemini 3.8 Flash
Google · Released September 2, 2026
Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.
Gemini 3.8 Flash facts and comparisons
The short answer
GPT-6 Luna costs far less: its rates are 7.5x below Gemini 3.8 Flash's, and a month of example agentic coding sessions comes to $11.55 against $78.38. The two sit in different tiers, since OpenAI pitches Luna for focused, high-volume work while Google pitches Gemini 3.8 Flash for long-horizon software engineering and autonomous agents. Gemini 3.8 Flash is on introductory rates through December 31, 2026, after which its listed prices double.
Choose GPT-6 Luna if
- Cost comes first: $0.10 input and $0.50 output per million, against $0.75 and $3.75.
- Your tasks are narrow and repeatable, the kind the Codex docs recommend Luna for.
- You need long responses: Luna writes up to 128K tokens against Gemini 3.8 Flash's 65.5K.
- Your team uses the Codex app on Free or Go plans, which include Luna.
Choose Gemini 3.8 Flash if
- Your work is long-horizon software engineering or multi-file refactoring, which Google names as Gemini 3.8 Flash's focus.
- You run Gemini CLI with a Gemini API key or Vertex AI, and its default auto model already uses Gemini 3.8 Flash as the Flash half.
- You want to start free: the Gemini API free tier covers its input, output, and caching.
- Your agents read untrusted content, and Google's claim of better robustness against prompt injection matters to you.
Side by side
Specs and prices
| Fact | GPT-6 Luna | Gemini 3.8 Flash |
|---|---|---|
| Maker | OpenAI | |
| API model id | gpt-6-luna | gemini-3.8-flash |
| Released | September 22, 2026 | September 2, 2026 |
| Status | Current | Current |
| Context window | 1.05M tokens | 1.05M tokens |
| Max output | 128K tokens | 65.5K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $0.10 | $0.75 |
| Cache hit, per 1M | $0.01 | $0.075 |
| Cache write, per 1M | $0.125 | $0.75 (same as input) |
| Output, per 1M tokens | $0.50 | $3.75 |
| Runs in | Codex, OpenCode, OpenRouter, and GitHub Copilot | Cursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. GPT-6 Luna: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | GPT-6 Luna | Gemini 3.8 Flash |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.11 | $0.71 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.02 | $0.15 |
| Output-heavy generation, 30K input, 80K output | $0.04 | $0.32 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $11.55 | $78.38 |
| Where the session’s cost goes | ||
| Cache writes | $0.05 | $0.30 |
| Cache reads | $0.02 | $0.15 |
| Uncached input | $0.01 | $0.08 |
| Output | $0.03 | $0.19 |
- caching saves on the session with GPT-6 Luna (61%)
- $0.17
- caching saves on the session with Gemini 3.8 Flash (66%)
- $1.35
Two different tiers, a 6.8x monthly gap
GPT-6 Luna is the low-cost end of the GPT-6 family. OpenAI calls it "our most efficient model for focused, high-volume tasks." Gemini 3.8 Flash is Google's newest Flash, described as "our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows." Comparing them sets OpenAI's budget tier against Google's main Flash tier.
The price difference follows. Luna charges $0.10 per million input tokens, $0.01 per million cache hits, and $0.50 per million output tokens. Gemini 3.8 Flash charges $0.75, $0.075, and $3.75, each 7.5x higher. The large one-off review costs $0.02 against $0.15.
Over 110 example sessions a month, Luna costs $11.55 and Gemini 3.8 Flash $78.38, a 6.8x gap and $66.83 apart. The monthly ratio is lower than the per-token ratio because cache writes are closer: OpenAI charges 1.25x input for a write, $0.125 per million on Luna, while Google bills written tokens as ordinary input at $0.75. That write gap is 6x rather than 7.5x.
What January 1, 2027 does to the gap
Gemini 3.8 Flash's current rates are introductory and run through December 31, 2026. From January 1, 2027, Google lists $1.50 per million input tokens, $0.15 per million cached tokens, and $7.50 per million output tokens, twice today's prices. Luna's price notes list no promotion, so on current information the gap between the two roughly doubles in the new year.
The session figures in this post use today's rates. If you are sizing a budget that runs into 2027, redo the Gemini side at the higher rates before comparing.
Thinking defaults, output limits, and caching
Both default to a medium setting: Luna's reasoning effort and Gemini 3.8 Flash's thinking level. Gemini CLI sends high for Gemini 3.8 Flash, and Codex lets Luna go up to max effort. Higher settings write more tokens, and output is 26% of Gemini 3.8 Flash's session cost here and 27% of Luna's.
Their context windows match at 1.05M tokens. Luna writes up to 128K tokens per response, twice Gemini 3.8 Flash's 65.5K. Luna's rates rise for any request over 272K input tokens: 2x for input and cache and 1.5x for output, for the whole request.
Caching saves 66% of the uncached session cost on Gemini 3.8 Flash, $1.35, and 61% on Luna, $0.17. Google's implicit caching is automatic but not assured to hit, and its explicit caching adds a storage charge of $0.50 to $1 per million tokens per hour on Flash models. OpenAI's cached prefixes stay reusable for at least 30 minutes after their last use.
Prompt caching
How each maker bills cached tokens
OpenAI
Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.
From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.
On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.
Source: OpenAI: Prompt caching
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what GPT-6 Luna and Gemini 3.8 Flash really cost you.
everyaitoken reads your Codex, OpenCode, OpenRouter, Gemini CLI, and Cursor history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
How much cheaper is GPT-6 Luna than Gemini 3.8 Flash?
7.5x per token on input, output, and cache hits, and 6.8x over a month of example agentic sessions, $11.55 against $78.38. The gap is set to widen when Gemini 3.8 Flash's introductory rates end.
Is there a free way to use Gemini 3.8 Flash?
Google's free tier on the Gemini API includes its input, output, and caching. Paid use costs $0.75 input and $3.75 output per million tokens until December 31, 2026.
Which of the two is aimed at complex coding?
Google positions Gemini 3.8 Flash for long-horizon software engineering, multi-file refactoring, and autonomous agents. OpenAI positions GPT-6 Luna for focused, repeatable, high-volume tasks, and the Codex docs recommend GPT-6 Sol for complex coding.
Where can I see both models' costs together?
EveryToken reads your local Codex and Gemini CLI history, prices each request at API rates, and shows cost and cache savings per model side by side.
Sources
- OpenAI: API pricing
- OpenAI docs: GPT-6 Luna
- OpenAI: Introducing GPT-6 Sol and Luna
- OpenAI: API changelog
- Codex docs: Models
- OpenRouter: GPT-6 Luna
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Google: Gemini API pricing
- Google docs: Gemini 3.8 Flash
- Google: Gemini models
- Google: Latest Gemini model
- Google DeepMind: Gemini 3.8 Flash model card
- Google: Gemini 3.8 Flash
- Gemini CLI source: model configuration
- Cursor docs: Gemini 3.8 Flash
- OpenRouter: Gemini 3.8 Flash
- OpenAI: Prompt caching
- Google: Context caching