Skip to content

Model comparison

GPT-5.6 Terra vs Gemini 3.8 Flash: rates, caching, status

GPT-5.6 Terra costs 3.1x as much as Gemini 3.8 Flash on a cached coding session, $2.20 against $0.71, and its listed 2027 rates stay below Terra's.

· Prices as of September 28, 2026

  • GPT-5.6 Terra

    OpenAI · Released July 9, 2026 · Previous generation

    The middle tier of GPT-5.6, balancing intelligence and cost, which roughly corresponds to earlier GPT-5 mini models.

    GPT-5.6 Terra facts and comparisons
  • Gemini 3.8 Flash

    Google · Released September 2, 2026

    Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.

    Gemini 3.8 Flash facts and comparisons

The short answer

Gemini 3.8 Flash is the cheaper model, with the example agentic coding session at $0.71 against $2.20 on GPT-5.6 Terra and output-heavy generation at $0.32 against $1.02. Its introductory rates end December 31, 2026, but the listed rates from January 1, 2027 ($1.50 input, $0.15 cached, $7.50 output) still sit below Terra's $2, $0.20, and $12. GPT-5.6 Terra is a previous model that Codex suggests replacing with GPT-6 Sol, so it mainly suits existing OpenAI setups.

Choose GPT-5.6 Terra if

  • Your code already calls gpt-5.6-terra, and it stays available in the API.
  • You need responses longer than 65.5K tokens: Terra writes up to 128K.
  • You work in Codex and want OpenAI's explicit cache breakpoints, with cached prefixes reusable for at least 30 minutes.
  • You value OpenAI's claim that GPT-5.6 infers the user's goal and intended level of work from context.

Choose Gemini 3.8 Flash if

  • You want the lower cost: $0.75 input and $3.75 output per million against $2 and $12, with the gap holding into 2027.
  • You use Gemini CLI, where Gemini 3.8 Flash is the Flash half of the default auto model for API key and Vertex AI users.
  • You want a current model that Google aims at long-horizon software engineering and multi-file refactoring.
  • You want to try it on the Gemini API free tier, which covers its input, output, and caching.

Side by side

Specs and prices

FactGPT-5.6 TerraGemini 3.8 Flash
MakerOpenAIGoogle
API model idgpt-5.6-terragemini-3.8-flash
ReleasedJuly 9, 2026September 2, 2026
StatusPrevious generationCurrent
Context window1.05M tokens1.05M tokens
Max output128K tokens65.5K tokens
Open weightsNoNo
Input, per 1M tokens$2$0.75
Cache hit, per 1M$0.20$0.075
Cache write, per 1M$2.50$0.75 (same as input)
Output, per 1M tokens$12$3.75
Runs inCodex, Cursor, OpenCode, OpenRouter, and GitHub CopilotCursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. GPT-5.6 Terra: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadGPT-5.6 TerraGemini 3.8 Flash
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$2.20$0.71
Large one-off review, 150K input with no cache hits, 10K output$0.42$0.15
Output-heavy generation, 30K input, 80K output$1.02$0.32
A month of sessions, 110 sessions: 5 a day, 22 working days$242.00$78.38
Where the session’s cost goes
Cache writes$1.00$0.30
Cache reads$0.40$0.15
Uncached input$0.20$0.08
Output$0.60$0.19
caching saves on the session with GPT-5.6 Terra (61%)
$3.40
caching saves on the session with Gemini 3.8 Flash (66%)
$1.35

How big is the gap, and where does it come from?

GPT-5.6 Terra charges $2 per million input tokens, $0.20 per million cache hits, $2.50 per million cache writes, and $12 per million output tokens. Gemini 3.8 Flash charges $0.75 for input, $0.075 for hits, and $3.75 for output, and bills cache writes as ordinary input at $0.75.

That puts Terra at 2.7x on input and hits, 3.2x on output, and 3.3x on writes, since OpenAI adds a 1.25x premium to writes from GPT-5.6 on and Google adds none. The large one-off review costs $0.42 against $0.15, and the output-heavy generation $1.02 against $0.32.

In the example agentic session, cache writes are the biggest line on both: $1.00 on Terra and $0.30 on Gemini, 45% and 42% of each total. Output adds another $0.41 of difference. The session comes to $2.20 against $0.71, and 110 sessions a month to $242.00 against $78.38, $163.62 apart.

Does the 2027 price rise change the answer?

Google's introductory pricing for Gemini 3.8 Flash runs through December 31, 2026. From January 1, 2027, Google lists $1.50 per million input tokens, $0.15 per million cached tokens, and $7.50 per million output tokens, double today's rates.

Even then, each rate stays below Terra's, and because Google bills cache writes at the input rate, the cached session would still cost less on Gemini 3.8 Flash than on Terra. The gap would narrow, not flip. Terra's own rates carry no promotion in its price notes; OpenAI cut them by 20% on July 30, 2026.

Status, limits, and how each maker pitches its model

Terra is a previous-generation model. OpenAI describes it as the "GPT-5.6 model that balances intelligence and cost," roughly where earlier GPT-5 mini models sat, and Codex now suggests moving to GPT-6 Sol. It stays in the API and in Codex, Cursor, OpenRouter, OpenCode, and GitHub Copilot.

Gemini 3.8 Flash is current, released on September 2, 2026. Google calls it "our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows" and claims fewer failed agent loops. Google defaults it to medium thinking, and Gemini CLI requests high. It runs in Gemini CLI, Cursor, OpenRouter, OpenCode, and GitHub Copilot.

Both have 1.05M-token context windows. Terra writes up to 128K tokens and Gemini 3.8 Flash 65.5K. Terra's rates rise for requests over 272K input tokens, to 2x for input and cache and 1.5x for output across the whole request. Google's explicit caching, when you use it instead of implicit caching, adds a storage charge of $0.50 to $1 per million tokens per hour on Flash models.

Prompt caching

How each maker bills cached tokens

OpenAI

Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.

From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.

On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.

Source: OpenAI: Prompt caching

Google

Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.

Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.

On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.

Source: Google: Context caching

Your own numbers

See what GPT-5.6 Terra and Gemini 3.8 Flash really cost you.

everyaitoken reads your Codex, Cursor, OpenCode, OpenRouter, and Gemini CLI history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is Gemini 3.8 Flash cheaper than GPT-5.6 Terra?

Yes, on every rate and workload. The example cached session costs $0.71 against $2.20, a 3.1x gap, and output-heavy generation $0.32 against $1.02.

What replaces GPT-5.6 Terra?

Codex suggests moving to GPT-6 Sol. Terra stays available in the OpenAI API, and no retirement date is listed for it.

Which one can write longer outputs?

GPT-5.6 Terra, at up to 128K tokens per response against 65.5K for Gemini 3.8 Flash. Context is level at 1.05M tokens on each.

How do I compare them on my own history?

EveryToken reads local Codex, Cursor, and Gemini CLI history, prices every request at API rates, and shows the cost and cache savings for each model. The figures are estimates at published rates.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash

    Gemini 3.8 Flash costs about half of Gemini 3.5 Flash on introductory rates. What changes on January 1, 2027, and which one Gemini CLI picks for you.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash-Lite

    Google recommends Gemini 3.8 Flash and Gemini 3.5 Flash-Lite together for new projects. Where each fits, the 2.1x session gap, and what changes in 2027.

  • Gemini 3.8 Flash vs Gemini 3.1 Pro Preview

    Gemini CLI's auto model uses both Gemini 3.8 Flash and Gemini 3.1 Pro Preview. What each costs, why Flash rates double in 2027, and the Pro's preview status.

  • Gemini 3.8 Flash vs Gemini 3.7 Flash

    Gemini 3.8 Flash and Gemini 3.7 Flash cost the same, and both rates double on January 1, 2027. What differs is Gemini CLI, Cursor, and Google's own pitch.

  • GPT-5.6 Sol vs GPT-5.6 Terra

    GPT-5.6 Sol costs about twice GPT-5.6 Terra, on promotional rates. How Terra's output price narrows the gap and why Codex points both to GPT-6 Sol.

  • GPT-5.6 Terra vs GPT-5.6 Luna

    GPT-5.6 Terra costs 10x GPT-5.6 Luna on every rate after OpenAI's July 30 price cuts. What each tier is for, and why Codex now sends them different ways.