Skip to content

Model comparison

Claude Sonnet 5 vs GPT-6 Astra: cost of OpenAI's top model

GPT-6 Astra charges 5x Claude Sonnet 5's rates on every kind of token. One cached coding session costs $10.50 against $2.40, a 4.4x gap. Here is why.

· Prices as of September 28, 2026

  • Claude Sonnet 5

    Anthropic · Released June 30, 2026

    Anthropic's balance of speed and intelligence, and a drop-in upgrade for Claude Sonnet 4.6. Anthropic pitches it as close to Claude Opus 4.8 at a lower price.

    Claude Sonnet 5 facts and comparisons
  • GPT-6 Astra

    OpenAI · Released September 3, 2026

    OpenAI's top GPT-6 model and its most expensive, aimed at the hardest long-running work that spans many tools, including coding.

    GPT-6 Astra facts and comparisons

The short answer

Claude Sonnet 5 is the far cheaper model: every GPT-6 Astra rate is 5x higher, from $10 against $2 for input to $50 against $10 for output. On the example agentic coding session the gap narrows to 4.4x, $10.50 against $2.40, because Anthropic's 1-hour cache writes cost 2x input. OpenAI reserves GPT-6 Astra for the hardest long-running work and makes it the default in Codex CLI's bundled model list, so pick it for that tier of job and Sonnet 5 for everyday agentic coding.

Choose Claude Sonnet 5 if

  • Cost per token matters to you: Sonnet 5's input, output, cache-write, and cache-hit rates are each a fifth of GPT-6 Astra's.
  • You run many sessions a day: 110 example sessions come to $264.00 on Sonnet 5 and $1,155.00 on GPT-6 Astra at API rates.
  • You work in Claude Code, which offers Claude Sonnet 5.
  • Anthropic's claim of performance close to Claude Opus 4.8 at lower prices fits the work you do.

Choose GPT-6 Astra if

  • You use Codex CLI, where GPT-6 Astra is the default in the bundled model list of version 0.158.0 and starts at low reasoning effort.
  • Your jobs are the long, end-to-end kind that span many tools, which is what OpenAI built GPT-6 Astra for.
  • You judge cost per finished task rather than per token, and OpenAI's claim that Astra reaches stronger results with substantially fewer output tokens holds on your work.

Side by side

Specs and prices

FactClaude Sonnet 5GPT-6 Astra
MakerAnthropicOpenAI
API model idclaude-sonnet-5gpt-6-astra
ReleasedJune 30, 2026September 3, 2026
StatusCurrentCurrent
Context window1M tokens1.05M tokens
Max output128K tokens128K tokens
Open weightsNoNo
Input, per 1M tokens$2$10
Cache hit, per 1M$0.20$1
Cache write, per 1M$2.50 (5-minute), $4 (1-hour)$12.50
Output, per 1M tokens$10$50
Runs inClaude Code, Cursor, OpenCode, OpenRouter, and GitHub CopilotCodex, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker (Claude Sonnet 5: September 26, 2026; GPT-6 Astra: September 28, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Claude Sonnet 5: The full 1M context window is billed at standard rates. The launch price became the standard price on August 10, 2026, and a planned increase was cancelled. GPT-6 Astra: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadClaude Sonnet 5GPT-6 Astra
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$2.40$10.50
Large one-off review, 150K input with no cache hits, 10K output$0.40$2.00
Output-heavy generation, 30K input, 80K output$0.86$4.30
A month of sessions, 110 sessions: 5 a day, 22 working days$264.00$1,155.00
Where the session’s cost goes
Cache writes$1.30$5.00
Cache reads$0.40$2.00
Uncached input$0.20$1.00
Output$0.50$2.50
caching saves on the session with Claude Sonnet 5 (56%)
$3.10
caching saves on the session with GPT-6 Astra (62%)
$17.00

How much more does GPT-6 Astra cost than Claude Sonnet 5?

Five times as much on every published rate. GPT-6 Astra charges $10 per million input tokens, $1 per million cache hits, $12.50 per million cache writes, and $50 per million output tokens. Claude Sonnet 5 charges $2 for input, $0.20 for hits, $2.50 for writes, and $10 for output. That makes Astra 400% more expensive, or Sonnet 5 80% cheaper, depending on which way you read it.

Uncached work shows the full ratio. The large one-off review costs $2.00 on Astra and $0.40 on Sonnet 5, and the output-heavy generation costs $4.30 against $0.86, a $3.44 difference on a single job.

The agentic session is the one place the ratio slips, to 4.4x: $10.50 against $2.40. Reads, fresh input, and output are exactly a fifth on Sonnet 5 ($0.40, $0.20, and $0.50 against $2.00, $1.00, and $2.50). Cache writes are not, because half of Sonnet 5's writes in this session use Anthropic's 1-hour cache at $4 per million, while Astra charges its single write price for all of them. Writes come to $1.30 on Sonnet 5 and $5.00 on Astra.

Can fewer output tokens close a 5x gap?

OpenAI's case for Astra is cost per task, not per token. It says Astra delivers stronger results with substantially fewer output tokens in several of its evaluations, for a lower estimated cost per task, and that it stays coherent on long tasks better than GPT-5.6 Sol and earlier models. Those are OpenAI's claims, and how they play out depends on the work.

The session's cost structure limits how much output efficiency can do. On Astra, output is 24% of the session's cost. Cache writes are 48% and cache reads 19%, and those are driven by how much context the tool resends, not by how much the model writes. A model that writes less trims the smaller share of the session.

Effort settings pull the other way. Codex CLI starts Astra at low reasoning effort, while Sonnet 5 defaults to high effort with adaptive thinking. Tokenizers also differ between the two makers. The per-token ratio is fixed, but the per-task ratio is something to measure on your own history.

Where each model runs, and the limits that apply

Astra is available in the Codex app, CLI, and IDE extension, but not in Codex cloud. It is also offered on OpenRouter, OpenCode, and GitHub Copilot. Sonnet 5 runs in Claude Code, Cursor, OpenRouter, OpenCode, and GitHub Copilot, and Claude Code's sonnet alias resolves to it on the Anthropic API.

Their output ceilings match at 128K tokens. Astra's window is 1.05M tokens with up to 922K of input, and Sonnet 5's is 1M. Astra's price notes add a long-context rule: requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request. At Astra's rates that premium is large in dollar terms.

Caching saves more in dollars on the dearer model. On the example session it saves $17.00 on Astra, or 62% of the uncached cost, against $3.10 and 56% on Sonnet 5. If you run Astra at all, keeping its cache warm is where most of the savings are.

Prompt caching

How each maker bills cached tokens

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

OpenAI

Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.

From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.

On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.

Source: OpenAI: Prompt caching

Your own numbers

See what Claude Sonnet 5 and GPT-6 Astra really cost you.

everyaitoken reads your Claude Code, Cursor, OpenCode, OpenRouter, and Codex history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is GPT-6 Astra five times the price of Claude Sonnet 5?

On the rate card, yes: input, output, cache writes, and cache hits all carry a 5x price. On the example agentic session the ratio is 4.4x, $10.50 against $2.40, because half of Sonnet 5's cache writes use Anthropic's 1-hour rate. A month of those sessions differs by $891.00.

Is GPT-6 Astra the default in Codex CLI?

GPT-6 Astra is the default in the bundled model list of Codex CLI version 0.158.0, where it starts at low reasoning effort. You can pick another model per session, and Astra is not available in Codex cloud.

Does GPT-6 Astra have a bigger context window than Claude Sonnet 5?

Slightly: 1.05M tokens against 1M, with up to 922K of input on Astra. Both cap output at 128K. On Astra, any request over 272K input tokens is billed at higher long-context rates.

How do I measure what Sonnet 5 and Astra cost me?

EveryToken reads your local Claude Code and Codex history, prices every request at the maker's API rates, and splits the total by model with cache savings shown. It is a $9 one-time macOS app, and its figures are API-equivalent estimates rather than what a subscription charges.

  • Claude Fable 5.1 vs Claude Sonnet 5

    Claude Fable 5.1 lists at 5x the rates of Claude Sonnet 5, but cache hits cost $0.25 against $0.20. Why sessions cost 4.4x, and who Anthropic aims Fable 5.1 at.

  • Claude Fable 5.1 vs GPT-6 Astra

    Claude Fable 5.1 and GPT-6 Astra share $10 input and $50 output prices, and a cached coding session costs $10.50 on each. Caching decides which way it tips.

  • Claude Opus 5.5 vs Claude Sonnet 5

    Claude Opus 5.5 costs twice as much per token as Claude Sonnet 5, yet cache hits cost $0.20 on both. What that means for agentic coding sessions.

  • Claude Opus 5.5 vs GPT-6 Astra

    Claude Code defaults to Claude Opus 5.5 and Codex CLI to GPT-6 Astra. Astra costs 2.5x more per token and 2.4x more on a cached coding session. Here is why.

  • Claude Sonnet 5 vs Claude Haiku 4.5

    Claude Haiku 4.5 costs half as much as Claude Sonnet 5, with a 200K context window and a retirement date ahead. When the cheaper model fits a coding workflow.

  • Claude Sonnet 5 vs Claude Sonnet 4.6

    Claude Sonnet 5 costs 33% less per token than Claude Sonnet 4.6, but its tokenizer counts up to 1.35x as many tokens. What the upgrade saves in practice.