Model comparison
GPT-6 Sol vs Gemini 3.8 Flash: 3x now, less in 2027
Gemini 3.8 Flash costs a third of GPT-6 Sol on a cached coding session, at introductory rates that end December 31, 2026. How the gap changes after that.
· Prices as of September 28, 2026
GPT-6 Sol
OpenAI · Released September 22, 2026
The mid-priced GPT-6 model, which OpenAI pitches for complex coding and agent workflows and which the Codex docs recommend for complex coding.
GPT-6 Sol facts and comparisonsGemini 3.8 Flash
Google · Released September 2, 2026
Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.
Gemini 3.8 Flash facts and comparisons
The short answer
Gemini 3.8 Flash is cheaper: the example agentic coding session costs $0.71 on it and $2.10 on GPT-6 Sol, a 3x gap, and 110 sessions a month come to $78.38 against $231.00. The gap shrinks on January 1, 2027, when Flash's introductory rates end and its prices double to $1.50 input and $7.50 output per million tokens. Choose Sol for Codex and outputs up to 128K tokens, and Flash for high-volume agent loops in Gemini CLI.
Choose GPT-6 Sol if
- Codex is your main tool, and its docs point complex coding to GPT-6 Sol.
- Your responses run long: Sol writes up to 128K tokens, Flash up to 65.5K.
- You want explicit cache control, with up to four breakpoints and a cached prefix that stays reusable for at least 30 minutes after its last use.
Choose Gemini 3.8 Flash if
- Price leads: Flash's input, output, and cache hits cost 63% less than Sol's today.
- Gemini CLI is your tool, signed in with a Gemini API key or through Vertex AI, where Flash already runs as half of the default auto model.
- You want to prototype on the free tier, which covers Flash's input, output, and caching.
- Google's claims fit your agents: long-horizon software engineering, complex multi-file refactoring, and deterministic tool execution.
Side by side
Specs and prices
| Fact | GPT-6 Sol | Gemini 3.8 Flash |
|---|---|---|
| Maker | OpenAI | |
| API model id | gpt-6-sol | gemini-3.8-flash |
| Released | September 22, 2026 | September 2, 2026 |
| Status | Current | Current |
| Context window | 1.05M tokens | 1.05M tokens |
| Max output | 128K tokens | 65.5K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $2 | $0.75 |
| Cache hit, per 1M | $0.20 | $0.075 |
| Cache write, per 1M | $2.50 | $0.75 (same as input) |
| Output, per 1M tokens | $10 | $3.75 |
| Runs in | Codex, OpenCode, OpenRouter, and GitHub Copilot | Cursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. GPT-6 Sol: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | GPT-6 Sol | Gemini 3.8 Flash |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $2.10 | $0.71 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.40 | $0.15 |
| Output-heavy generation, 30K input, 80K output | $0.86 | $0.32 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $231.00 | $78.38 |
| Where the session’s cost goes | ||
| Cache writes | $1.00 | $0.30 |
| Cache reads | $0.40 | $0.15 |
| Uncached input | $0.20 | $0.08 |
| Output | $0.50 | $0.19 |
- caching saves on the session with GPT-6 Sol (62%)
- $3.40
- caching saves on the session with Gemini 3.8 Flash (66%)
- $1.35
Why a 2.7x price gap becomes 3x on the session
GPT-6 Sol charges $2 input, $0.20 per cache hit, and $10 output per million tokens. Gemini 3.8 Flash charges $0.75, $0.075, and $3.75, so all three rates sit 2.7x apart. The uncached review, $0.40 against $0.15, and output-heavy generation, $0.86 against $0.32, keep that ratio.
Cache writes widen it on the agentic session. OpenAI bills a Sol write at 1.25x input, $2.50 per million, while Google bills Flash's written tokens as ordinary input at $0.75, a 3.3x gap. Writes cost $1.00 on Sol and $0.30 on Flash, $0.70 of the $1.39 session difference, and the session lands at $2.10 against $0.71.
What Gemini 3.8 Flash costs after the introductory period
Google lists Flash's current prices as introductory through December 31, 2026. For 2027 it has published $1.50 input, $0.15 cached, and $7.50 output per million tokens. At those rates Sol's input, cache hit, and output prices would be a third higher than Flash's, instead of well over double.
Written tokens on Flash would then cost $1.50, against Sol's $2.50 write price. The 3x session gap of today becomes a much smaller one in 2027, so a budget that spans the new year should use the higher Flash rates. Explicit Flash caches also carry a storage charge of $0.50 to $1 per million tokens per hour, which these estimates leave out.
OpenAI's middle model against Google's top Flash
OpenAI places Sol in the middle of its GPT-6 line, above GPT-6 Luna and below GPT-6 Astra, and says it is "Built to power complex coding and agentic workflows." Google calls Gemini 3.8 Flash "Our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows." Both are pitched at agentic coding below their makers' most expensive options.
Both default to medium reasoning, though Gemini CLI sends Flash a high thinking level. Sol accepts up to 922K input tokens of its 1.05M window and switches to long-context rates above 272K, at 2x for input and cache and 1.5x for output. Flash's window is also 1.05M, but it writes at most 65.5K tokens per response against Sol's 128K.
The tables price the same tokens on both models, while in practice reasoning settings change how many tokens each writes. EveryToken reads your Codex and Gemini CLI history on a Mac and shows each model's API-equivalent cost and cache savings from your real sessions.
Prompt caching
How each maker bills cached tokens
OpenAI
Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.
From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.
On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.
Source: OpenAI: Prompt caching
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what GPT-6 Sol and Gemini 3.8 Flash really cost you.
everyaitoken reads your Codex, OpenCode, OpenRouter, Gemini CLI, and Cursor history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
How much cheaper is Gemini 3.8 Flash than GPT-6 Sol?
At current rates it is 63% cheaper on input, output, and cache hits, and 70% cheaper on cache writes. The example agentic session costs $0.71 on Flash against $2.10 on Sol, a $1.39 difference.
What will Gemini 3.8 Flash cost in 2027?
Google has set $1.50 input, $0.15 cached, and $7.50 output per million tokens from January 1, 2027. That is twice today's introductory rates.
Are both models in GitHub Copilot?
Yes, GitHub Copilot lists both GPT-6 Sol and Gemini 3.8 Flash. Both also run in OpenCode and OpenRouter, while Sol runs in Codex and Flash in Gemini CLI.
Do both discount cache hits the same way?
Yes, each charges 10% of input for a hit. The difference is in writes: OpenAI adds a 1.25x premium on Sol, and Google bills Flash's written tokens as ordinary input.
Sources
- OpenAI: API pricing
- OpenAI docs: GPT-6 Sol
- OpenAI: Introducing GPT-6 Sol and Luna
- OpenAI: API changelog
- Codex docs: Models
- OpenRouter: GPT-6 Sol
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Google: Gemini API pricing
- Google docs: Gemini 3.8 Flash
- Google: Gemini models
- Google: Latest Gemini model
- Google DeepMind: Gemini 3.8 Flash model card
- Google: Gemini 3.8 Flash
- Gemini CLI source: model configuration
- Cursor docs: Gemini 3.8 Flash
- OpenRouter: Gemini 3.8 Flash
- OpenAI: Prompt caching
- Google: Context caching