Model comparison
GPT-6 Astra vs Gemini 3.1 Pro Preview: top models priced
GPT-6 Astra costs 5.3x as much as Gemini 3.1 Pro Preview on a cached coding session. How OpenAI's write premium, long-context tiers, and output caps compare.
· Prices as of September 28, 2026
GPT-6 Astra
OpenAI · Released September 3, 2026
OpenAI's top GPT-6 model and its most expensive, aimed at the hardest long-running work that spans many tools, including coding.
GPT-6 Astra facts and comparisonsGemini 3.1 Pro Preview
Google · Released February 19, 2026 · Preview
Google's only current Pro model, still in preview, positioned for deep reasoning and agentic coding.
Gemini 3.1 Pro Preview facts and comparisons
The short answer
Gemini 3.1 Pro Preview costs far less: the example agentic coding session is $2.00 on it and $10.50 on GPT-6 Astra, a 5.3x gap, with input prices 5x apart. GPT-6 Astra, which OpenAI calls its most capable model, is also its most expensive and Codex CLI's default, and writes up to 128K tokens per response. Gemini 3.1 Pro Preview is Google's current Pro model, much cheaper but still a preview with a 65.5K output cap, and the Pro half of Gemini CLI's default.
Choose GPT-6 Astra if
- You work in Codex, where the CLI's bundled model list sets Astra as the default, or in GitHub Copilot.
- You need up to 128K output tokens per response, against 65.5K on Gemini 3.1 Pro Preview.
- You want a current release rather than a preview whose successor, Gemini 3.5 Pro, has been announced.
- OpenAI's claims describe your work: the hardest end-to-end tasks across many tools, with stronger results from fewer output tokens by OpenAI's account.
Choose Gemini 3.1 Pro Preview if
- Cost matters most: 110 sessions a month cost $220.00, against $1,155.00 on Astra.
- You work in Gemini CLI, where it is the Pro half of the default auto model.
- You write the cache heavily, since Google adds no write premium while OpenAI charges 1.25x input.
Side by side
Specs and prices
| Fact | GPT-6 Astra | Gemini 3.1 Pro Preview |
|---|---|---|
| Maker | OpenAI | |
| API model id | gpt-6-astra | gemini-3.1-pro-preview |
| Released | September 3, 2026 | February 19, 2026 |
| Status | Current | Preview |
| Context window | 1.05M tokens | 1.05M tokens |
| Max output | 128K tokens | 65.5K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $10 | $2 |
| Cache hit, per 1M | $1 | $0.20 |
| Cache write, per 1M | $12.50 | $2 (same as input) |
| Output, per 1M tokens | $50 | $12 |
| Runs in | Codex, OpenCode, OpenRouter, and GitHub Copilot | Cursor, Gemini CLI, OpenCode, and OpenRouter |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. GPT-6 Astra: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request. Gemini 3.1 Pro Preview: Prompts over 200K input tokens cost $4 input, $0.40 cached, and $18 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | GPT-6 Astra | Gemini 3.1 Pro Preview |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $10.50 | $2.00 |
| Large one-off review, 150K input with no cache hits, 10K output | $2.00 | $0.42 |
| Output-heavy generation, 30K input, 80K output | $4.30 | $1.02 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $1,155.00 | $220.00 |
| Where the session’s cost goes | ||
| Cache writes | $5.00 | $0.80 |
| Cache reads | $2.00 | $0.40 |
| Uncached input | $1.00 | $0.20 |
| Output | $2.50 | $0.60 |
- caching saves on the session with GPT-6 Astra (62%)
- $17.00
- caching saves on the session with Gemini 3.1 Pro Preview (64%)
- $3.60
Same cache discount, different write rules
Both makers charge 10% of input for a cache hit, so the $1 per million of GPT-6 Astra and the $0.20 of Gemini 3.1 Pro Preview keep the 5x input ratio. The difference is in writing the cache. From GPT-5.6 on, OpenAI charges 1.25x input for a write, $12.50 per million on Astra, while Google has no separate write price and bills written tokens at the $2 input rate.
That puts writes 6.3x apart per token and pushes the session gap above the list-price gap. Cache writes cost $5.00 on Astra and $0.80 on Gemini, and the session totals $10.50 against $2.00. Over 110 sessions a month that is $1,155.00 against $220.00, a $935.00 difference in API-equivalent terms.
Without cache writes the gap narrows: 4.8x on the uncached review, $2.00 against $0.42, and 4.2x on output-heavy generation, $4.30 against $1.02, where the output gap is smaller than the input gap.
Two long-context tiers with different thresholds
Both models advertise a 1.05M context window, and both charge more for long prompts, starting at different points. Google raises Gemini 3.1 Pro Preview to $4 input, $0.40 cached, and $18 output for prompts over 200K input tokens. OpenAI bills any Astra request over 272K input tokens at 2x for input and cache and 1.5x for output, for the whole request, and accepts at most 922K input tokens.
So a prompt between 200K and 272K tokens pays Gemini's higher tier and Astra's standard one, and a prompt above 272K pays more on both. The example session stays under 200K per request, so the cost table uses standard rates for each.
General availability against preview
Astra was released on September 3, 2026, and runs in the Codex app, CLI, and IDE extension, though not in Codex cloud, as well as in OpenCode, OpenRouter, and GitHub Copilot. Codex starts it at low reasoning effort. OpenAI calls it "Our most capable model, built for the hardest end-to-end work" and says it reached stronger results with substantially fewer output tokens in several evaluations.
Gemini 3.1 Pro Preview has been a preview since February 19, 2026. Google has announced Gemini 3.5 Pro, not yet released, and GitHub Copilot retired Gemini 3.1 Pro Preview on September 1, 2026. Its thinking cannot be turned off and defaults to high, and thinking settings change how many tokens a task uses on either model.
Because both makers' claims concern token counts that the cost table holds fixed, the useful test is your own work. EveryToken reads Codex and Gemini CLI history on a Mac and prices each request at API rates, with cache savings shown per model.
Prompt caching
How each maker bills cached tokens
OpenAI
Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.
From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.
On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.
Source: OpenAI: Prompt caching
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what GPT-6 Astra and Gemini 3.1 Pro Preview really cost you.
everyaitoken reads your Codex, OpenCode, OpenRouter, Gemini CLI, and Cursor history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
How much more does GPT-6 Astra cost than Gemini 3.1 Pro Preview?
It costs 5x more for input and cache hits, 4.2x more for output, and 6.3x more for cache writes. On the example agentic session, Astra costs $10.50 and Gemini $2.00, a difference of $8.50.
Which has the bigger context window?
Both list 1.05M tokens. GPT-6 Astra accepts up to 922K of that as input. Gemini 3.1 Pro Preview's pricing changes above 200K input tokens, and Astra's above 272K.
Why does a GPT-6 Astra cache write cost more than its input?
From GPT-5.6 on, OpenAI charges 1.25x the uncached input price to write the cache, and a hit then costs 0.1x. You can mark up to four explicit cache breakpoints, and a cached prefix stays reusable for at least 30 minutes after its last use.
Is Gemini 3.1 Pro Preview available in GitHub Copilot?
Not anymore. GitHub Copilot retired it on September 1, 2026, while GPT-6 Astra is available there. Gemini 3.1 Pro Preview remains in Gemini CLI, Cursor, OpenCode, and OpenRouter.
Sources
- OpenAI: API pricing
- OpenAI docs: GPT-6 Astra
- OpenAI: Using the latest model
- OpenAI: API changelog
- Codex docs: Models
- Codex CLI: bundled model catalog
- OpenRouter: GPT-6 Astra
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Google: Gemini API pricing
- Google docs: Gemini 3.1 Pro Preview
- Google: Gemini models
- Google: Gemini 3.1 Pro
- Google DeepMind: Gemini
- Gemini CLI source: model configuration
- Cursor docs: Gemini 3.1 Pro
- OpenRouter: Gemini 3.1 Pro Preview
- OpenAI: Prompt caching
- Google: Context caching