Skip to content

Model comparison

GPT-6 Luna vs Gemini 3.5 Flash-Lite: budget model costs

GPT-6 Luna lists a third of Gemini 3.5 Flash-Lite's input rate and a fifth of its output rate. A month of cached coding sessions: $11.55 against $36.85.

· Prices as of September 28, 2026

  • GPT-6 Luna

    OpenAI · Released September 22, 2026

    The cheapest GPT-6 model, pitched for focused, high-volume, repeatable work, including narrower coding tasks.

    GPT-6 Luna facts and comparisons
  • Gemini 3.5 Flash-Lite

    Google · Released July 21, 2026

    Google's current low-cost, low-latency tier for high-volume and subagent work, recommended with Gemini 3.8 Flash for new projects.

    Gemini 3.5 Flash-Lite facts and comparisons

The short answer

GPT-6 Luna is cheaper on every rate, at $0.10 input and $0.50 output per million against $0.30 and $2.50 for Gemini 3.5 Flash-Lite, and a month of example agentic coding sessions costs $11.55 against $36.85. The gap is widest on output-heavy work, since Flash-Lite's output rate is 5x Luna's. Luna fits Codex and GitHub Copilot, and Flash-Lite fits Gemini CLI and high-throughput sub-agents.

Choose GPT-6 Luna if

  • You want the lower rates: a third of Flash-Lite's input and cache-hit prices and a fifth of its output price.
  • Your jobs write a lot: Luna can return up to 128K tokens in one response against Flash-Lite's 65.5K.
  • You work in Codex, whose docs recommend Luna for focused, repeatable tasks, or on a Free or Go plan that gets it in the Codex app.
  • You use GitHub Copilot, which offers GPT-6 Luna and not Gemini 3.5 Flash-Lite.

Choose Gemini 3.5 Flash-Lite if

  • You use Gemini CLI, where Flash-Lite is the model behind the flash-lite alias.
  • Raw output speed counts, and Google cites about 350 output tokens per second.
  • Your agent needs computer use, which Google builds in as a tool for Flash-Lite.
  • You run subagent tasks or document parsing, the uses Google's docs name for Flash-Lite.

Side by side

Specs and prices

FactGPT-6 LunaGemini 3.5 Flash-Lite
MakerOpenAIGoogle
API model idgpt-6-lunagemini-3.5-flash-lite
ReleasedSeptember 22, 2026July 21, 2026
StatusCurrentCurrent
Context window1.05M tokens1.05M tokens
Max output128K tokens65.5K tokens
Open weightsNoNo
Input, per 1M tokens$0.10$0.30
Cache hit, per 1M$0.01$0.03
Cache write, per 1M$0.125$0.30 (same as input)
Output, per 1M tokens$0.50$2.50
Runs inCodex, OpenCode, OpenRouter, and GitHub CopilotGemini CLI, OpenCode, and OpenRouter

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. GPT-6 Luna: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadGPT-6 LunaGemini 3.5 Flash-Lite
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$0.11$0.34
Large one-off review, 150K input with no cache hits, 10K output$0.02$0.07
Output-heavy generation, 30K input, 80K output$0.04$0.21
A month of sessions, 110 sessions: 5 a day, 22 working days$11.55$36.85
Where the session’s cost goes
Cache writes$0.05$0.12
Cache reads$0.02$0.06
Uncached input$0.01$0.03
Output$0.03$0.13
caching saves on the session with GPT-6 Luna (61%)
$0.17
caching saves on the session with Gemini 3.5 Flash-Lite (61%)
$0.54

Where the gap comes from: output first, then input

GPT-6 Luna charges $0.10 per million input tokens, $0.01 per million cache hits, and $0.50 per million output tokens. Gemini 3.5 Flash-Lite charges $0.30 for input, $0.03 for hits, and $2.50 for output. Flash-Lite costs 3x as much on input and hits and 5x as much on output.

At these prices single jobs round to a few cents, so ratios on one job move around. The output-heavy generation costs $0.04 on Luna and $0.21 on Flash-Lite. The large one-off review costs $0.02 against $0.07. The monthly projection is steadier: 110 example sessions cost $11.55 on Luna and $36.85 on Flash-Lite, a 3.2x gap and $25.30 apart.

Why the session gap is smaller than the output gap

Cache writes are the line where Flash-Lite comes closest. OpenAI charges 1.25x input for a cache write on its newer models, $0.125 per million on Luna. Google publishes no write price and bills written tokens as ordinary input, $0.30 per million. So the write gap is 2.4x, well under the 5x output gap.

In the example session, which writes 400K tokens and reads 2M from the cache, writes cost $0.05 on Luna and $0.12 on Flash-Lite, and output $0.03 against $0.13. Output is 38% of Flash-Lite's session cost and 27% of Luna's. The session totals are $0.11 and $0.34.

Caching saves the same share on both, 61% of what the session would cost uncached: $0.17 on Luna and $0.54 on Flash-Lite. Google's implicit caching applies discounts automatically but does not promise a hit, so real sessions can save less than this.

Defaults, limits, and where each one runs

The two start from different thinking defaults. Luna's reasoning effort defaults to medium and goes up to max in Codex. Flash-Lite starts at a minimal thinking level. Higher settings mean more output tokens, and output is the rate where these models differ most, so a like-for-like comparison should match effort as closely as the settings allow.

Both have 1.05M-token context windows. Luna's notes add a long-context rule: any request over 272K input tokens costs 2x for input and cache and 1.5x for output, for the whole request. Luna writes up to 128K tokens per response, and Flash-Lite 65.5K.

OpenAI calls Luna "our most efficient model for focused, high-volume tasks." Google positions Flash-Lite as its low-cost, low-latency tier for high-volume and sub-agent work, recommended alongside Gemini 3.8 Flash for new projects. Luna runs in Codex (not Codex cloud), OpenRouter, OpenCode, and GitHub Copilot. Flash-Lite runs in Gemini CLI, OpenRouter, and OpenCode, and neither Cursor nor Copilot offers it.

Prompt caching

How each maker bills cached tokens

OpenAI

Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.

From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.

On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.

Source: OpenAI: Prompt caching

Google

Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.

Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.

On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.

Source: Google: Context caching

Your own numbers

See what GPT-6 Luna and Gemini 3.5 Flash-Lite really cost you.

everyaitoken reads your Codex, OpenCode, OpenRouter, and Gemini CLI history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Which is cheaper, GPT-6 Luna or Gemini 3.5 Flash-Lite?

GPT-6 Luna, on every rate. A month of example agentic sessions costs $11.55 on Luna and $36.85 on Flash-Lite, and output-heavy work is where Luna's lead is largest.

Do both models have the same context window?

Yes, 1.05M tokens each. GPT-6 Luna writes up to 128K tokens of output and Gemini 3.5 Flash-Lite up to 65.5K, and Luna's rates rise on requests over 272K input tokens.

Does Gemini 3.5 Flash-Lite charge for cache writes?

Google publishes no separate write price, so written tokens cost ordinary input, $0.30 per million. GPT-6 Luna charges 1.25x input for a write, $0.125 per million, which is still lower in absolute terms.

How do I compare them on my own usage?

EveryToken reads Codex and Gemini CLI history on your Mac, prices every request at API rates, and shows each model's cost with cache savings. At budget prices like these, the monthly total is the number worth watching.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash-Lite

    Google recommends Gemini 3.8 Flash and Gemini 3.5 Flash-Lite together for new projects. Where each fits, the 2.1x session gap, and what changes in 2027.

  • GPT-5.6 Luna vs Gemini 3.5 Flash-Lite

    Gemini 3.5 Flash-Lite costs 1.5x as much as GPT-5.6 Luna on a cached coding session, $0.34 against $0.22, and about twice as much on output-heavy work.

  • GPT-6 Astra vs GPT-6 Luna

    GPT-6 Astra and GPT-6 Luna sit at opposite ends of OpenAI's lineup, 100x apart per token. What each is for, and where GPT-6 Sol fits between them.

  • GPT-6 Luna vs Gemini 3.8 Flash

    Gemini 3.8 Flash costs 7.5x as much as GPT-6 Luna per token, and its introductory rates end December 31, 2026. A month of sessions: $11.55 vs $78.38.

  • GPT-6 Sol vs GPT-6 Luna

    GPT-6 Luna costs a twentieth of GPT-6 Sol per token. What OpenAI and the Codex docs say each tier is for, and what the gap means for coding sessions.

  • Gemini 3.5 Flash-Lite vs Gemini 3.1 Flash-Lite

    Gemini 3.1 Flash-Lite shuts down on May 7, 2027, and Gemini 3.5 Flash-Lite replaces it at higher rates. What the move costs, mostly on output.