Model comparison
GPT-6 Luna vs Gemini 3.5 Flash-Lite: budget model costs
GPT-6 Luna lists a third of Gemini 3.5 Flash-Lite's input rate and a fifth of its output rate. A month of cached coding sessions: $11.55 against $36.85.
· Prices as of September 28, 2026
GPT-6 Luna
OpenAI · Released September 22, 2026
The cheapest GPT-6 model, pitched for focused, high-volume, repeatable work, including narrower coding tasks.
GPT-6 Luna facts and comparisonsGemini 3.5 Flash-Lite
Google · Released July 21, 2026
Google's current low-cost, low-latency tier for high-volume and subagent work, recommended with Gemini 3.8 Flash for new projects.
Gemini 3.5 Flash-Lite facts and comparisons
The short answer
GPT-6 Luna is cheaper on every rate, at $0.10 input and $0.50 output per million against $0.30 and $2.50 for Gemini 3.5 Flash-Lite, and a month of example agentic coding sessions costs $11.55 against $36.85. The gap is widest on output-heavy work, since Flash-Lite's output rate is 5x Luna's. Luna fits Codex and GitHub Copilot, and Flash-Lite fits Gemini CLI and high-throughput sub-agents.
Choose GPT-6 Luna if
- You want the lower rates: a third of Flash-Lite's input and cache-hit prices and a fifth of its output price.
- Your jobs write a lot: Luna can return up to 128K tokens in one response against Flash-Lite's 65.5K.
- You work in Codex, whose docs recommend Luna for focused, repeatable tasks, or on a Free or Go plan that gets it in the Codex app.
- You use GitHub Copilot, which offers GPT-6 Luna and not Gemini 3.5 Flash-Lite.
Choose Gemini 3.5 Flash-Lite if
- You use Gemini CLI, where Flash-Lite is the model behind the flash-lite alias.
- Raw output speed counts, and Google cites about 350 output tokens per second.
- Your agent needs computer use, which Google builds in as a tool for Flash-Lite.
- You run subagent tasks or document parsing, the uses Google's docs name for Flash-Lite.
Side by side
Specs and prices
| Fact | GPT-6 Luna | Gemini 3.5 Flash-Lite |
|---|---|---|
| Maker | OpenAI | |
| API model id | gpt-6-luna | gemini-3.5-flash-lite |
| Released | September 22, 2026 | July 21, 2026 |
| Status | Current | Current |
| Context window | 1.05M tokens | 1.05M tokens |
| Max output | 128K tokens | 65.5K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $0.10 | $0.30 |
| Cache hit, per 1M | $0.01 | $0.03 |
| Cache write, per 1M | $0.125 | $0.30 (same as input) |
| Output, per 1M tokens | $0.50 | $2.50 |
| Runs in | Codex, OpenCode, OpenRouter, and GitHub Copilot | Gemini CLI, OpenCode, and OpenRouter |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. GPT-6 Luna: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | GPT-6 Luna | Gemini 3.5 Flash-Lite |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.11 | $0.34 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.02 | $0.07 |
| Output-heavy generation, 30K input, 80K output | $0.04 | $0.21 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $11.55 | $36.85 |
| Where the session’s cost goes | ||
| Cache writes | $0.05 | $0.12 |
| Cache reads | $0.02 | $0.06 |
| Uncached input | $0.01 | $0.03 |
| Output | $0.03 | $0.13 |
- caching saves on the session with GPT-6 Luna (61%)
- $0.17
- caching saves on the session with Gemini 3.5 Flash-Lite (61%)
- $0.54
Where the gap comes from: output first, then input
GPT-6 Luna charges $0.10 per million input tokens, $0.01 per million cache hits, and $0.50 per million output tokens. Gemini 3.5 Flash-Lite charges $0.30 for input, $0.03 for hits, and $2.50 for output. Flash-Lite costs 3x as much on input and hits and 5x as much on output.
At these prices single jobs round to a few cents, so ratios on one job move around. The output-heavy generation costs $0.04 on Luna and $0.21 on Flash-Lite. The large one-off review costs $0.02 against $0.07. The monthly projection is steadier: 110 example sessions cost $11.55 on Luna and $36.85 on Flash-Lite, a 3.2x gap and $25.30 apart.
Why the session gap is smaller than the output gap
Cache writes are the line where Flash-Lite comes closest. OpenAI charges 1.25x input for a cache write on its newer models, $0.125 per million on Luna. Google publishes no write price and bills written tokens as ordinary input, $0.30 per million. So the write gap is 2.4x, well under the 5x output gap.
In the example session, which writes 400K tokens and reads 2M from the cache, writes cost $0.05 on Luna and $0.12 on Flash-Lite, and output $0.03 against $0.13. Output is 38% of Flash-Lite's session cost and 27% of Luna's. The session totals are $0.11 and $0.34.
Caching saves the same share on both, 61% of what the session would cost uncached: $0.17 on Luna and $0.54 on Flash-Lite. Google's implicit caching applies discounts automatically but does not promise a hit, so real sessions can save less than this.
Defaults, limits, and where each one runs
The two start from different thinking defaults. Luna's reasoning effort defaults to medium and goes up to max in Codex. Flash-Lite starts at a minimal thinking level. Higher settings mean more output tokens, and output is the rate where these models differ most, so a like-for-like comparison should match effort as closely as the settings allow.
Both have 1.05M-token context windows. Luna's notes add a long-context rule: any request over 272K input tokens costs 2x for input and cache and 1.5x for output, for the whole request. Luna writes up to 128K tokens per response, and Flash-Lite 65.5K.
OpenAI calls Luna "our most efficient model for focused, high-volume tasks." Google positions Flash-Lite as its low-cost, low-latency tier for high-volume and sub-agent work, recommended alongside Gemini 3.8 Flash for new projects. Luna runs in Codex (not Codex cloud), OpenRouter, OpenCode, and GitHub Copilot. Flash-Lite runs in Gemini CLI, OpenRouter, and OpenCode, and neither Cursor nor Copilot offers it.
Prompt caching
How each maker bills cached tokens
OpenAI
Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.
From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.
On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.
Source: OpenAI: Prompt caching
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what GPT-6 Luna and Gemini 3.5 Flash-Lite really cost you.
everyaitoken reads your Codex, OpenCode, OpenRouter, and Gemini CLI history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Which is cheaper, GPT-6 Luna or Gemini 3.5 Flash-Lite?
GPT-6 Luna, on every rate. A month of example agentic sessions costs $11.55 on Luna and $36.85 on Flash-Lite, and output-heavy work is where Luna's lead is largest.
Do both models have the same context window?
Yes, 1.05M tokens each. GPT-6 Luna writes up to 128K tokens of output and Gemini 3.5 Flash-Lite up to 65.5K, and Luna's rates rise on requests over 272K input tokens.
Does Gemini 3.5 Flash-Lite charge for cache writes?
Google publishes no separate write price, so written tokens cost ordinary input, $0.30 per million. GPT-6 Luna charges 1.25x input for a write, $0.125 per million, which is still lower in absolute terms.
How do I compare them on my own usage?
EveryToken reads Codex and Gemini CLI history on your Mac, prices every request at API rates, and shows each model's cost with cache savings. At budget prices like these, the monthly total is the number worth watching.
Sources
- OpenAI: API pricing
- OpenAI docs: GPT-6 Luna
- OpenAI: Introducing GPT-6 Sol and Luna
- OpenAI: API changelog
- Codex docs: Models
- OpenRouter: GPT-6 Luna
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Google: Gemini API pricing
- Google docs: Gemini 3.5 Flash-Lite
- Google: Gemini models
- Google: Gemini 3.6 Flash and 3.5 Flash-Lite
- Gemini CLI source: model configuration
- OpenRouter: Gemini 3.5 Flash-Lite
- OpenAI: Prompt caching
- Google: Context caching