Skip to content

OpenAI

GPT-5.6 Terra: price, context window, and caching

The middle tier of GPT-5.6, balancing intelligence and cost, which roughly corresponds to earlier GPT-5 mini models.

Released July 9, 2026 · Prices as of September 28, 2026

In OpenAI’s words

“GPT-5.6 model that balances intelligence and cost”

OpenAI docs: GPT-5.6 Terra

What OpenAI says it’s good at

  • Inferring the user's underlying goal and intended level of work from context, like the rest of GPT-5.6 Source
  • Token efficiency and better frontend design, like the rest of GPT-5.6 Source

Facts

Specs and prices

FactGPT-5.6 Terra
MakerOpenAI
API model idgpt-5.6-terra
ReleasedJuly 9, 2026
StatusPrevious generation
Context window1.05M tokens
Max output128K tokens
Open weightsNo
Input, per 1M tokens$2
Cache hit, per 1M$0.20
Cache write, per 1M$2.50
Output, per 1M tokens$12
Runs inCodex, Cursor, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by the maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. GPT-5.6 Terra: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.

Good to know

  • Codex suggests moving to GPT-6 Sol. It stays available in the API.
  • OpenAI cut its price by 20% on July 30, 2026.

Cost

What typical work costs

Example token counts at GPT-5.6 Terra’s published rates. On the agentic session, caching saves $3.40 against billing every token as ordinary input.

Example workload costs for GPT-5.6 Terra
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$2.20
Large one-off review, 150K input with no cache hits, 10K output$0.42
Output-heavy generation, 30K input, 80K output$1.02
A month of sessions, 110 sessions: 5 a day, 22 working days$242.00

Prompt caching

How OpenAI bills cached tokens

OpenAI

Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.

From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.

On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.

Source: OpenAI: Prompt caching

GPT-5.6 Terra compared

  • Claude Haiku 4.5 vs GPT-5.6 Terra

    Claude Haiku 4.5 costs about half of GPT-5.6 Terra per token, $1.20 against $2.20 per coding session. Terra offers a 1.05M context window and 128K output.

  • Claude Sonnet 5 vs GPT-5.6 Terra

    Claude Sonnet 5 and GPT-5.6 Terra share a $2 input rate. Terra costs less on cached sessions, Sonnet 5 on output-heavy work. The numbers, line by line.

  • GPT-5.6 Sol vs GPT-5.6 Terra

    GPT-5.6 Sol costs about twice GPT-5.6 Terra, on promotional rates. How Terra's output price narrows the gap and why Codex points both to GPT-6 Sol.

  • GPT-5.6 Terra vs Gemini 3.8 Flash

    GPT-5.6 Terra costs 3.1x as much as Gemini 3.8 Flash on a cached coding session, $2.20 against $0.71, and its listed 2027 rates stay below Terra's.

  • GPT-5.6 Terra vs GPT-5.6 Luna

    GPT-5.6 Terra costs 10x GPT-5.6 Luna on every rate after OpenAI's July 30 price cuts. What each tier is for, and why Codex now sends them different ways.

  • GPT-6 Sol vs GPT-5.6 Terra

    GPT-6 Sol matches GPT-5.6 Terra's input and cache prices and charges less for output. What that means for coding sessions and the move Codex suggests.

Your own numbers

See what GPT-5.6 Terra really costs you.

everyaitoken reads your Codex, Cursor, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math