Skip to content

Model comparison

GPT-6 Sol vs Gemini 3.8 Flash: 3x now, less in 2027

Gemini 3.8 Flash costs a third of GPT-6 Sol on a cached coding session, at introductory rates that end December 31, 2026. How the gap changes after that.

· Prices as of September 28, 2026

  • GPT-6 Sol

    OpenAI · Released September 22, 2026

    The mid-priced GPT-6 model, which OpenAI pitches for complex coding and agent workflows and which the Codex docs recommend for complex coding.

    GPT-6 Sol facts and comparisons
  • Gemini 3.8 Flash

    Google · Released September 2, 2026

    Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.

    Gemini 3.8 Flash facts and comparisons

The short answer

Gemini 3.8 Flash is cheaper: the example agentic coding session costs $0.71 on it and $2.10 on GPT-6 Sol, a 3x gap, and 110 sessions a month come to $78.38 against $231.00. The gap shrinks on January 1, 2027, when Flash's introductory rates end and its prices double to $1.50 input and $7.50 output per million tokens. Choose Sol for Codex and outputs up to 128K tokens, and Flash for high-volume agent loops in Gemini CLI.

Choose GPT-6 Sol if

  • Codex is your main tool, and its docs point complex coding to GPT-6 Sol.
  • Your responses run long: Sol writes up to 128K tokens, Flash up to 65.5K.
  • You want explicit cache control, with up to four breakpoints and a cached prefix that stays reusable for at least 30 minutes after its last use.

Choose Gemini 3.8 Flash if

  • Price leads: Flash's input, output, and cache hits cost 63% less than Sol's today.
  • Gemini CLI is your tool, signed in with a Gemini API key or through Vertex AI, where Flash already runs as half of the default auto model.
  • You want to prototype on the free tier, which covers Flash's input, output, and caching.
  • Google's claims fit your agents: long-horizon software engineering, complex multi-file refactoring, and deterministic tool execution.

Side by side

Specs and prices

FactGPT-6 SolGemini 3.8 Flash
MakerOpenAIGoogle
API model idgpt-6-solgemini-3.8-flash
ReleasedSeptember 22, 2026September 2, 2026
StatusCurrentCurrent
Context window1.05M tokens1.05M tokens
Max output128K tokens65.5K tokens
Open weightsNoNo
Input, per 1M tokens$2$0.75
Cache hit, per 1M$0.20$0.075
Cache write, per 1M$2.50$0.75 (same as input)
Output, per 1M tokens$10$3.75
Runs inCodex, OpenCode, OpenRouter, and GitHub CopilotCursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. GPT-6 Sol: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadGPT-6 SolGemini 3.8 Flash
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$2.10$0.71
Large one-off review, 150K input with no cache hits, 10K output$0.40$0.15
Output-heavy generation, 30K input, 80K output$0.86$0.32
A month of sessions, 110 sessions: 5 a day, 22 working days$231.00$78.38
Where the session’s cost goes
Cache writes$1.00$0.30
Cache reads$0.40$0.15
Uncached input$0.20$0.08
Output$0.50$0.19
caching saves on the session with GPT-6 Sol (62%)
$3.40
caching saves on the session with Gemini 3.8 Flash (66%)
$1.35

Why a 2.7x price gap becomes 3x on the session

GPT-6 Sol charges $2 input, $0.20 per cache hit, and $10 output per million tokens. Gemini 3.8 Flash charges $0.75, $0.075, and $3.75, so all three rates sit 2.7x apart. The uncached review, $0.40 against $0.15, and output-heavy generation, $0.86 against $0.32, keep that ratio.

Cache writes widen it on the agentic session. OpenAI bills a Sol write at 1.25x input, $2.50 per million, while Google bills Flash's written tokens as ordinary input at $0.75, a 3.3x gap. Writes cost $1.00 on Sol and $0.30 on Flash, $0.70 of the $1.39 session difference, and the session lands at $2.10 against $0.71.

What Gemini 3.8 Flash costs after the introductory period

Google lists Flash's current prices as introductory through December 31, 2026. For 2027 it has published $1.50 input, $0.15 cached, and $7.50 output per million tokens. At those rates Sol's input, cache hit, and output prices would be a third higher than Flash's, instead of well over double.

Written tokens on Flash would then cost $1.50, against Sol's $2.50 write price. The 3x session gap of today becomes a much smaller one in 2027, so a budget that spans the new year should use the higher Flash rates. Explicit Flash caches also carry a storage charge of $0.50 to $1 per million tokens per hour, which these estimates leave out.

OpenAI's middle model against Google's top Flash

OpenAI places Sol in the middle of its GPT-6 line, above GPT-6 Luna and below GPT-6 Astra, and says it is "Built to power complex coding and agentic workflows." Google calls Gemini 3.8 Flash "Our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows." Both are pitched at agentic coding below their makers' most expensive options.

Both default to medium reasoning, though Gemini CLI sends Flash a high thinking level. Sol accepts up to 922K input tokens of its 1.05M window and switches to long-context rates above 272K, at 2x for input and cache and 1.5x for output. Flash's window is also 1.05M, but it writes at most 65.5K tokens per response against Sol's 128K.

The tables price the same tokens on both models, while in practice reasoning settings change how many tokens each writes. EveryToken reads your Codex and Gemini CLI history on a Mac and shows each model's API-equivalent cost and cache savings from your real sessions.

Prompt caching

How each maker bills cached tokens

OpenAI

Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.

From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.

On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.

Source: OpenAI: Prompt caching

Google

Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.

Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.

On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.

Source: Google: Context caching

Your own numbers

See what GPT-6 Sol and Gemini 3.8 Flash really cost you.

everyaitoken reads your Codex, OpenCode, OpenRouter, Gemini CLI, and Cursor history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

How much cheaper is Gemini 3.8 Flash than GPT-6 Sol?

At current rates it is 63% cheaper on input, output, and cache hits, and 70% cheaper on cache writes. The example agentic session costs $0.71 on Flash against $2.10 on Sol, a $1.39 difference.

What will Gemini 3.8 Flash cost in 2027?

Google has set $1.50 input, $0.15 cached, and $7.50 output per million tokens from January 1, 2027. That is twice today's introductory rates.

Are both models in GitHub Copilot?

Yes, GitHub Copilot lists both GPT-6 Sol and Gemini 3.8 Flash. Both also run in OpenCode and OpenRouter, while Sol runs in Codex and Flash in Gemini CLI.

Do both discount cache hits the same way?

Yes, each charges 10% of input for a hit. The difference is in writes: OpenAI adds a 1.25x premium on Sol, and Google bills Flash's written tokens as ordinary input.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash

    Gemini 3.8 Flash costs about half of Gemini 3.5 Flash on introductory rates. What changes on January 1, 2027, and which one Gemini CLI picks for you.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash-Lite

    Google recommends Gemini 3.8 Flash and Gemini 3.5 Flash-Lite together for new projects. Where each fits, the 2.1x session gap, and what changes in 2027.

  • Gemini 3.8 Flash vs Gemini 3.1 Pro Preview

    Gemini CLI's auto model uses both Gemini 3.8 Flash and Gemini 3.1 Pro Preview. What each costs, why Flash rates double in 2027, and the Pro's preview status.

  • Gemini 3.8 Flash vs Gemini 3.7 Flash

    Gemini 3.8 Flash and Gemini 3.7 Flash cost the same, and both rates double on January 1, 2027. What differs is Gemini CLI, Cursor, and Google's own pitch.

  • GPT-5.6 Terra vs Gemini 3.8 Flash

    GPT-5.6 Terra costs 3.1x as much as Gemini 3.8 Flash on a cached coding session, $2.20 against $0.71, and its listed 2027 rates stay below Terra's.

  • GPT-6 Astra vs Gemini 3.8 Flash

    GPT-6 Astra and Gemini 3.8 Flash launched a day apart at opposite ends of the price range, 14.8x apart on a cached coding session. What each is built for.