Skip to content

Model comparison

GPT-6 Astra vs Gemini 3.1 Pro Preview: top models priced

GPT-6 Astra costs 5.3x as much as Gemini 3.1 Pro Preview on a cached coding session. How OpenAI's write premium, long-context tiers, and output caps compare.

· Prices as of September 28, 2026

  • GPT-6 Astra

    OpenAI · Released September 3, 2026

    OpenAI's top GPT-6 model and its most expensive, aimed at the hardest long-running work that spans many tools, including coding.

    GPT-6 Astra facts and comparisons
  • Gemini 3.1 Pro Preview

    Google · Released February 19, 2026 · Preview

    Google's only current Pro model, still in preview, positioned for deep reasoning and agentic coding.

    Gemini 3.1 Pro Preview facts and comparisons

The short answer

Gemini 3.1 Pro Preview costs far less: the example agentic coding session is $2.00 on it and $10.50 on GPT-6 Astra, a 5.3x gap, with input prices 5x apart. GPT-6 Astra, which OpenAI calls its most capable model, is also its most expensive and Codex CLI's default, and writes up to 128K tokens per response. Gemini 3.1 Pro Preview is Google's current Pro model, much cheaper but still a preview with a 65.5K output cap, and the Pro half of Gemini CLI's default.

Choose GPT-6 Astra if

  • You work in Codex, where the CLI's bundled model list sets Astra as the default, or in GitHub Copilot.
  • You need up to 128K output tokens per response, against 65.5K on Gemini 3.1 Pro Preview.
  • You want a current release rather than a preview whose successor, Gemini 3.5 Pro, has been announced.
  • OpenAI's claims describe your work: the hardest end-to-end tasks across many tools, with stronger results from fewer output tokens by OpenAI's account.

Choose Gemini 3.1 Pro Preview if

  • Cost matters most: 110 sessions a month cost $220.00, against $1,155.00 on Astra.
  • You work in Gemini CLI, where it is the Pro half of the default auto model.
  • You write the cache heavily, since Google adds no write premium while OpenAI charges 1.25x input.

Side by side

Specs and prices

FactGPT-6 AstraGemini 3.1 Pro Preview
MakerOpenAIGoogle
API model idgpt-6-astragemini-3.1-pro-preview
ReleasedSeptember 3, 2026February 19, 2026
StatusCurrentPreview
Context window1.05M tokens1.05M tokens
Max output128K tokens65.5K tokens
Open weightsNoNo
Input, per 1M tokens$10$2
Cache hit, per 1M$1$0.20
Cache write, per 1M$12.50$2 (same as input)
Output, per 1M tokens$50$12
Runs inCodex, OpenCode, OpenRouter, and GitHub CopilotCursor, Gemini CLI, OpenCode, and OpenRouter

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. GPT-6 Astra: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request. Gemini 3.1 Pro Preview: Prompts over 200K input tokens cost $4 input, $0.40 cached, and $18 output per million tokens.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadGPT-6 AstraGemini 3.1 Pro Preview
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$10.50$2.00
Large one-off review, 150K input with no cache hits, 10K output$2.00$0.42
Output-heavy generation, 30K input, 80K output$4.30$1.02
A month of sessions, 110 sessions: 5 a day, 22 working days$1,155.00$220.00
Where the session’s cost goes
Cache writes$5.00$0.80
Cache reads$2.00$0.40
Uncached input$1.00$0.20
Output$2.50$0.60
caching saves on the session with GPT-6 Astra (62%)
$17.00
caching saves on the session with Gemini 3.1 Pro Preview (64%)
$3.60

Same cache discount, different write rules

Both makers charge 10% of input for a cache hit, so the $1 per million of GPT-6 Astra and the $0.20 of Gemini 3.1 Pro Preview keep the 5x input ratio. The difference is in writing the cache. From GPT-5.6 on, OpenAI charges 1.25x input for a write, $12.50 per million on Astra, while Google has no separate write price and bills written tokens at the $2 input rate.

That puts writes 6.3x apart per token and pushes the session gap above the list-price gap. Cache writes cost $5.00 on Astra and $0.80 on Gemini, and the session totals $10.50 against $2.00. Over 110 sessions a month that is $1,155.00 against $220.00, a $935.00 difference in API-equivalent terms.

Without cache writes the gap narrows: 4.8x on the uncached review, $2.00 against $0.42, and 4.2x on output-heavy generation, $4.30 against $1.02, where the output gap is smaller than the input gap.

Two long-context tiers with different thresholds

Both models advertise a 1.05M context window, and both charge more for long prompts, starting at different points. Google raises Gemini 3.1 Pro Preview to $4 input, $0.40 cached, and $18 output for prompts over 200K input tokens. OpenAI bills any Astra request over 272K input tokens at 2x for input and cache and 1.5x for output, for the whole request, and accepts at most 922K input tokens.

So a prompt between 200K and 272K tokens pays Gemini's higher tier and Astra's standard one, and a prompt above 272K pays more on both. The example session stays under 200K per request, so the cost table uses standard rates for each.

General availability against preview

Astra was released on September 3, 2026, and runs in the Codex app, CLI, and IDE extension, though not in Codex cloud, as well as in OpenCode, OpenRouter, and GitHub Copilot. Codex starts it at low reasoning effort. OpenAI calls it "Our most capable model, built for the hardest end-to-end work" and says it reached stronger results with substantially fewer output tokens in several evaluations.

Gemini 3.1 Pro Preview has been a preview since February 19, 2026. Google has announced Gemini 3.5 Pro, not yet released, and GitHub Copilot retired Gemini 3.1 Pro Preview on September 1, 2026. Its thinking cannot be turned off and defaults to high, and thinking settings change how many tokens a task uses on either model.

Because both makers' claims concern token counts that the cost table holds fixed, the useful test is your own work. EveryToken reads Codex and Gemini CLI history on a Mac and prices each request at API rates, with cache savings shown per model.

Prompt caching

How each maker bills cached tokens

OpenAI

Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.

From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.

On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.

Source: OpenAI: Prompt caching

Google

Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.

Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.

On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.

Source: Google: Context caching

Your own numbers

See what GPT-6 Astra and Gemini 3.1 Pro Preview really cost you.

everyaitoken reads your Codex, OpenCode, OpenRouter, Gemini CLI, and Cursor history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

How much more does GPT-6 Astra cost than Gemini 3.1 Pro Preview?

It costs 5x more for input and cache hits, 4.2x more for output, and 6.3x more for cache writes. On the example agentic session, Astra costs $10.50 and Gemini $2.00, a difference of $8.50.

Which has the bigger context window?

Both list 1.05M tokens. GPT-6 Astra accepts up to 922K of that as input. Gemini 3.1 Pro Preview's pricing changes above 200K input tokens, and Astra's above 272K.

Why does a GPT-6 Astra cache write cost more than its input?

From GPT-5.6 on, OpenAI charges 1.25x the uncached input price to write the cache, and a hit then costs 0.1x. You can mark up to four explicit cache breakpoints, and a cached prefix stays reusable for at least 30 minutes after its last use.

Is Gemini 3.1 Pro Preview available in GitHub Copilot?

Not anymore. GitHub Copilot retired it on September 1, 2026, while GPT-6 Astra is available there. Gemini 3.1 Pro Preview remains in Gemini CLI, Cursor, OpenCode, and OpenRouter.

  • Gemini 3.1 Pro Preview vs Gemini 2.5 Pro

    Gemini 3.1 Pro Preview costs 60% more per input token than Gemini 2.5 Pro and is still a preview, while 2.5 Pro now limits new access. How to choose.

  • Gemini 3.8 Flash vs Gemini 3.1 Pro Preview

    Gemini CLI's auto model uses both Gemini 3.8 Flash and Gemini 3.1 Pro Preview. What each costs, why Flash rates double in 2027, and the Pro's preview status.

  • GPT-5.3-Codex vs Gemini 3.1 Pro Preview

    GPT-5.3-Codex and Gemini 3.1 Pro Preview cost within 4% of each other on a cached coding session. Context size, output limits, and access set them apart.

  • GPT-5.5 vs Gemini 3.1 Pro Preview

    GPT-5.5 costs 2.5x as much as Gemini 3.1 Pro Preview on every rate and every workload. Where they differ instead: output limits, long prompts, and access.

  • GPT-5.6 Sol vs Gemini 3.1 Pro Preview

    GPT-5.6 Sol is on promotional rates and Gemini 3.1 Pro Preview is still in preview. On a cached coding session, Gemini costs $2.00 against $4.20.

  • GPT-6 Astra vs Gemini 3.8 Flash

    GPT-6 Astra and Gemini 3.8 Flash launched a day apart at opposite ends of the price range, 14.8x apart on a cached coding session. What each is built for.