Skip to content

Model comparison

GPT-6 Astra vs Gemini 3.8 Flash: top tier or Flash tier?

GPT-6 Astra and Gemini 3.8 Flash launched a day apart at opposite ends of the price range, 14.8x apart on a cached coding session. What each is built for.

· Prices as of September 28, 2026

  • GPT-6 Astra

    OpenAI · Released September 3, 2026

    OpenAI's top GPT-6 model and its most expensive, aimed at the hardest long-running work that spans many tools, including coding.

    GPT-6 Astra facts and comparisons
  • Gemini 3.8 Flash

    Google · Released September 2, 2026

    Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.

    Gemini 3.8 Flash facts and comparisons

The short answer

Gemini 3.8 Flash is the budget choice by a wide margin: the example agentic coding session costs $0.71 on it and $10.50 on GPT-6 Astra, a 14.8x gap. Astra fits the hardest end-to-end work in Codex, where it is the CLI's default, and Flash fits high-volume agent loops in Gemini CLI. Flash's rates are introductory through December 31, 2026 and double after that, so budget with both price sets.

Choose GPT-6 Astra if

  • OpenAI's description matches your hardest tasks: long-running work across many tools, where it says Astra stays coherent better than GPT-5.6 Sol and earlier models.
  • You need up to 128K output tokens per response, against Flash's 65.5K.
  • You work in Codex and want to keep the CLI's default model.

Choose Gemini 3.8 Flash if

  • You run agents at volume: 110 sessions a month cost $78.38 on Flash against $1,155.00 on Astra.
  • You work in Gemini CLI with an API key or Vertex AI, where Flash is half of the default auto model.
  • You want to prototype on the Gemini API free tier, which covers Flash's input, output, and caching.
  • Prompt injection worries you, and Google claims better robustness against it for Flash.

Side by side

Specs and prices

FactGPT-6 AstraGemini 3.8 Flash
MakerOpenAIGoogle
API model idgpt-6-astragemini-3.8-flash
ReleasedSeptember 3, 2026September 2, 2026
StatusCurrentCurrent
Context window1.05M tokens1.05M tokens
Max output128K tokens65.5K tokens
Open weightsNoNo
Input, per 1M tokens$10$0.75
Cache hit, per 1M$1$0.075
Cache write, per 1M$12.50$0.75 (same as input)
Output, per 1M tokens$50$3.75
Runs inCodex, OpenCode, OpenRouter, and GitHub CopilotCursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. GPT-6 Astra: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadGPT-6 AstraGemini 3.8 Flash
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$10.50$0.71
Large one-off review, 150K input with no cache hits, 10K output$2.00$0.15
Output-heavy generation, 30K input, 80K output$4.30$0.32
A month of sessions, 110 sessions: 5 a day, 22 working days$1,155.00$78.38
Where the session’s cost goes
Cache writes$5.00$0.30
Cache reads$2.00$0.15
Uncached input$1.00$0.08
Output$2.50$0.19
caching saves on the session with GPT-6 Astra (62%)
$17.00
caching saves on the session with Gemini 3.8 Flash (66%)
$1.35

Both discount cache hits by 90%, so why is the gap 14.8x?

On input, output, and cache hits, GPT-6 Astra costs 13.3x what Gemini 3.8 Flash does: $10 against $0.75, $50 against $3.75, and $1 against $0.075 per million tokens. Both makers price a hit at 10% of input, so caching on its own would leave the ratio where it is.

Cache writes change it. OpenAI charges 1.25x input to write the cache on GPT-5.6 and later models, $12.50 per million on Astra, while Google bills written tokens as ordinary input, $0.75. That is a 16.7x gap per written token. In the session, writes cost $5.00 on Astra and $0.30 on Flash, lifting the total gap to 14.8x, $10.50 against $0.71.

Over 110 sessions a month the example reaches $1,155.00 on Astra and $78.38 on Flash, a $1,076.62 difference. Both totals are API-equivalent estimates; ChatGPT and Gemini plans that include usage are priced differently.

The 2027 price step for Gemini 3.8 Flash

Flash's rates are introductory through December 31, 2026. On January 1, 2027 they move to $1.50 input, $0.15 cached, and $7.50 output per million tokens. Every Flash figure on this page doubles at that point, and the gap to Astra roughly halves.

The two models launched a day apart, Flash on September 2 and Astra on September 3, 2026. Astra's price notes list no promotional period, only a long-context surcharge: above 272K input tokens, OpenAI bills the whole request at 2x for input and cache and 1.5x for output. Astra accepts at most 922K input tokens of its 1.05M window, and Flash lists a 1.05M window as well.

Output tokens per task: what the tables can't show

Both makers make efficiency claims rather than price claims. OpenAI says Astra reached stronger results with substantially fewer output tokens in several evaluations, for a lower estimated cost per task. Google says Flash handles multi-step planning and tool orchestration with fewer failed loops. Both claims concern how many tokens a task consumes, which the cost table holds fixed.

Default settings pull in different directions too. Codex starts Astra at low reasoning effort, while Flash defaults to medium thinking and Gemini CLI sends high. Flash also caps output at 65.5K tokens per response, against Astra's 128K. The way to settle it is to run the same tasks on both and compare real token counts, which EveryToken reads from your Codex and Gemini CLI history and prices at API rates.

Prompt caching

How each maker bills cached tokens

OpenAI

Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.

From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.

On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.

Source: OpenAI: Prompt caching

Google

Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.

Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.

On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.

Source: Google: Context caching

Your own numbers

See what GPT-6 Astra and Gemini 3.8 Flash really cost you.

everyaitoken reads your Codex, OpenCode, OpenRouter, Gemini CLI, and Cursor history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

How much cheaper is Gemini 3.8 Flash than GPT-6 Astra?

At current rates, 93% cheaper on input, output, and cache hits. The example agentic session costs $0.71 on Flash and $10.50 on Astra. From January 1, 2027, Flash's rates double.

Were GPT-6 Astra and Gemini 3.8 Flash released at the same time?

Almost. Google released Gemini 3.8 Flash on September 2, 2026, and OpenAI released GPT-6 Astra the next day.

Which one is the default in Codex and Gemini CLI?

Codex CLI version 0.158.0 ships GPT-6 Astra as the default in its bundled model list. Gemini 3.8 Flash is the Flash half of Gemini CLI's default auto model for Gemini API key and Vertex AI users.

Is GPT-6 Astra worth 14.8x the price?

That depends on whether its results on your hardest tasks save more than the difference. OpenAI's claim of fewer output tokens per task could narrow the gap on real work, but that claim is OpenAI's and has not been checked here. Compare both on a sample of your own tasks before committing.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash

    Gemini 3.8 Flash costs about half of Gemini 3.5 Flash on introductory rates. What changes on January 1, 2027, and which one Gemini CLI picks for you.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash-Lite

    Google recommends Gemini 3.8 Flash and Gemini 3.5 Flash-Lite together for new projects. Where each fits, the 2.1x session gap, and what changes in 2027.

  • Gemini 3.8 Flash vs Gemini 3.1 Pro Preview

    Gemini CLI's auto model uses both Gemini 3.8 Flash and Gemini 3.1 Pro Preview. What each costs, why Flash rates double in 2027, and the Pro's preview status.

  • Gemini 3.8 Flash vs Gemini 3.7 Flash

    Gemini 3.8 Flash and Gemini 3.7 Flash cost the same, and both rates double on January 1, 2027. What differs is Gemini CLI, Cursor, and Google's own pitch.

  • GPT-5.6 Terra vs Gemini 3.8 Flash

    GPT-5.6 Terra costs 3.1x as much as Gemini 3.8 Flash on a cached coding session, $2.20 against $0.71, and its listed 2027 rates stay below Terra's.

  • GPT-6 Astra vs Gemini 3.1 Pro Preview

    GPT-6 Astra costs 5.3x as much as Gemini 3.1 Pro Preview on a cached coding session. How OpenAI's write premium, long-context tiers, and output caps compare.