Skip to content

Model comparison

GPT-5.6 Luna vs Gemini 3.5 Flash-Lite: July's budget models

Gemini 3.5 Flash-Lite costs 1.5x as much as GPT-5.6 Luna on a cached coding session, $0.34 against $0.22, and about twice as much on output-heavy work.

· Prices as of September 28, 2026

  • GPT-5.6 Luna

    OpenAI · Released July 9, 2026 · Previous generation

    The low-cost GPT-5.6 tier for cost-sensitive, high-volume work, which roughly corresponds to earlier GPT-5 nano models.

    GPT-5.6 Luna facts and comparisons
  • Gemini 3.5 Flash-Lite

    Google · Released July 21, 2026

    Google's current low-cost, low-latency tier for high-volume and subagent work, recommended with Gemini 3.8 Flash for new projects.

    Gemini 3.5 Flash-Lite facts and comparisons

The short answer

GPT-5.6 Luna is the cheaper of the two, at $0.22 against $0.34 for the example agentic coding session and $0.10 against $0.21 for output-heavy generation. Gemini 3.5 Flash-Lite is Google's current low-cost tier, while GPT-5.6 Luna is a previous model that Codex suggests replacing with GPT-6 Luna. The tool you use settles most cases: Luna runs in Codex and Cursor, Flash-Lite in Gemini CLI.

Choose GPT-5.6 Luna if

  • You want the lower rates: $0.20 input and $1.20 output per million, against $0.30 and $2.50.
  • You work in Cursor or GitHub Copilot, which offer GPT-5.6 Luna and not Gemini 3.5 Flash-Lite.
  • Your jobs return long outputs: Luna writes up to 128K tokens per response against 65.5K.
  • Your code already depends on gpt-5.6-luna, which OpenAI keeps in the API.

Choose Gemini 3.5 Flash-Lite if

  • You want a current model: Google recommends Flash-Lite alongside Gemini 3.8 Flash for new projects.
  • You use Gemini CLI, where the flash-lite alias maps to it.
  • You want throughput, where Google cites about 350 output tokens per second, or built-in computer use.

Side by side

Specs and prices

FactGPT-5.6 LunaGemini 3.5 Flash-Lite
MakerOpenAIGoogle
API model idgpt-5.6-lunagemini-3.5-flash-lite
ReleasedJuly 9, 2026July 21, 2026
StatusPrevious generationCurrent
Context window1.05M tokens1.05M tokens
Max output128K tokens65.5K tokens
Open weightsNoNo
Input, per 1M tokens$0.20$0.30
Cache hit, per 1M$0.02$0.03
Cache write, per 1M$0.25$0.30 (same as input)
Output, per 1M tokens$1.20$2.50
Runs inCodex, Cursor, OpenCode, OpenRouter, and GitHub CopilotGemini CLI, OpenCode, and OpenRouter

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. GPT-5.6 Luna: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadGPT-5.6 LunaGemini 3.5 Flash-Lite
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$0.22$0.34
Large one-off review, 150K input with no cache hits, 10K output$0.04$0.07
Output-heavy generation, 30K input, 80K output$0.10$0.21
A month of sessions, 110 sessions: 5 a day, 22 working days$24.20$36.85
Where the session’s cost goes
Cache writes$0.10$0.12
Cache reads$0.04$0.06
Uncached input$0.02$0.03
Output$0.06$0.13
caching saves on the session with GPT-5.6 Luna (61%)
$0.34
caching saves on the session with Gemini 3.5 Flash-Lite (61%)
$0.54

How far apart are the prices?

Within about 2x on every rate. GPT-5.6 Luna charges $0.20 per million input tokens, $0.02 per million cache hits, $0.25 per million cache writes, and $1.20 per million output tokens. Gemini 3.5 Flash-Lite charges $0.30 for input, $0.03 for hits, and $2.50 for output, and bills cache writes as ordinary input at $0.30.

So Flash-Lite costs 50% more for input and cache hits, 20% more for cache writes, and 108% more for output. Luna's rates reflect an 80% price cut OpenAI made on July 30, 2026.

On the workloads, the large one-off review costs $0.04 on Luna and $0.07 on Flash-Lite, and the output-heavy generation $0.10 against $0.21. A month of 110 example sessions comes to $24.20 against $36.85, a $12.65 difference.

The output rate drives the difference

In the example agentic session, the cache-write line is nearly level: $0.10 on Luna and $0.12 on Flash-Lite. Luna's lower input price is mostly offset by OpenAI's 1.25x write premium, while Google adds none. Reads and fresh input differ by a cent or two.

Output is where the money goes. It costs $0.06 on Luna and $0.13 on Flash-Lite, $0.07 of the session's $0.12 gap, and it is 38% of Flash-Lite's session cost against 27% of Luna's. Any setting that makes both models write more pushes the comparison further toward Luna.

Caching saves 61% of the uncached session cost on both, $0.34 on Luna and $0.54 on Flash-Lite. OpenAI keeps cached prefixes reusable for at least 30 minutes on GPT-5.6 models and lets you mark up to four breakpoints. Google's implicit caching is automatic but a hit is not assured.

Launched the same month, now on different paths

Both launched in July 2026, Luna on July 9 and Flash-Lite on July 21. Since then their paths have split. Flash-Lite is Google's current low-cost, low-latency tier for high-volume and sub-agent work. GPT-5.6 Luna is now a previous model: Codex suggests moving to GPT-6 Luna, which lists lower rates, and GPT-5.6 Luna stays in the API.

OpenAI describes GPT-5.6 Luna as a "GPT-5.6 model optimized for cost-sensitive workloads," roughly where earlier GPT-5 nano models sat, and credits the GPT-5.6 family with reaching flagship-level performance using fewer output tokens. Google cites about 350 output tokens per second for Flash-Lite and names sub-agent tasks and document parsing as uses, and it sets Flash-Lite's thinking level to minimal by default.

Both have 1.05M-token context windows. Luna's rates rise to 2x for input and cache and 1.5x for output on any request over 272K input tokens.

Prompt caching

How each maker bills cached tokens

OpenAI

Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.

From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.

On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.

Source: OpenAI: Prompt caching

Google

Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.

Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.

On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.

Source: Google: Context caching

Your own numbers

See what GPT-5.6 Luna and Gemini 3.5 Flash-Lite really cost you.

everyaitoken reads your Codex, Cursor, OpenCode, OpenRouter, and Gemini CLI history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is GPT-5.6 Luna cheaper than Gemini 3.5 Flash-Lite?

Yes, on every rate. The example agentic session costs $0.22 against $0.34, and output-heavy generation $0.10 against $0.21.

Should I move from GPT-5.6 Luna to GPT-6 Luna?

Codex suggests it, and GPT-6 Luna lists lower rates than GPT-5.6 Luna. GPT-5.6 Luna stays in the API if existing code depends on it.

Which tools offer each model?

GPT-5.6 Luna is in Codex, Cursor, OpenRouter, OpenCode, and GitHub Copilot. Gemini 3.5 Flash-Lite is in Gemini CLI, OpenRouter, and OpenCode, and is not offered in Cursor or Copilot.

How can I compare their cost on my own work?

EveryToken reads local Codex, Cursor, Gemini CLI, OpenCode, and OpenRouter history, prices each request at API rates, and shows what caching saved per model. At budget rates like these, it helps to watch the monthly total rather than single requests.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash-Lite

    Google recommends Gemini 3.8 Flash and Gemini 3.5 Flash-Lite together for new projects. Where each fits, the 2.1x session gap, and what changes in 2027.

  • GPT-5.6 Terra vs GPT-5.6 Luna

    GPT-5.6 Terra costs 10x GPT-5.6 Luna on every rate after OpenAI's July 30 price cuts. What each tier is for, and why Codex now sends them different ways.

  • GPT-6 Luna vs Gemini 3.5 Flash-Lite

    GPT-6 Luna lists a third of Gemini 3.5 Flash-Lite's input rate and a fifth of its output rate. A month of cached coding sessions: $11.55 against $36.85.

  • Gemini 3.5 Flash-Lite vs Gemini 3.1 Flash-Lite

    Gemini 3.1 Flash-Lite shuts down on May 7, 2027, and Gemini 3.5 Flash-Lite replaces it at higher rates. What the move costs, mostly on output.

  • GPT-6 Luna vs GPT-5.6 Luna

    GPT-6 Luna halves the input price of GPT-5.6 Luna, which already had an 80% cut, and trims output further. What the move saves, and what Codex suggests.

  • Claude Haiku 4.5 vs Gemini 3.5 Flash-Lite

    Claude Haiku 4.5 and Gemini 3.5 Flash-Lite both target sub-agent work. Haiku 4.5 costs 3.5x as much per cached session, but only 2x as much per output token.