Model comparison
Claude Sonnet 4.6 or GPT-5.3-Codex? Context and caching
Claude Sonnet 4.6 and GPT-5.3-Codex launched 12 days apart in February 2026. One takes 1M tokens of context, the other costs 46% less per session. Tradeoffs.
· Prices as of September 28, 2026
Claude Sonnet 4.6
Anthropic · Released February 17, 2026 · Previous generation
The previous Sonnet, pitched at launch as near-Opus capability at the Sonnet price. Anthropic now recommends moving to Claude Sonnet 5.
Claude Sonnet 4.6 facts and comparisonsGPT-5.3-Codex
OpenAI · Released February 5, 2026 · Previous generation
The February 2026 Codex-tuned coding model, which combined GPT-5.2-Codex's coding with stronger reasoning.
GPT-5.3-Codex facts and comparisons
The short answer
GPT-5.3-Codex costs less on the example agentic coding session, $1.93 against $3.60 on Claude Sonnet 4.6, because its input is cheaper and OpenAI adds no charge for cache writes. Sonnet 4.6 accepts 1M tokens of context, while GPT-5.3-Codex takes up to 272K input tokens of its 400K window. Both are previous-generation models, and GPT-5.3-Codex is no longer selectable in Codex with ChatGPT sign-in.
Choose Claude Sonnet 4.6 if
- Your requests need more than 272K input tokens, which GPT-5.3-Codex won't accept.
- You work in Claude Code, where Sonnet 4.6 stays available as a legacy model.
- You want Anthropic's 1-hour cache for sessions with long gaps between turns.
Choose GPT-5.3-Codex if
- Your sessions are cache-heavy and fit in 272K input tokens, where GPT-5.3-Codex costs 46% less than Sonnet 4.6 on the example.
- You use Codex or another tool with an OpenAI API key, where GPT-5.3-Codex remains available.
- You want the model OpenAI tuned for Codex, which it says combined GPT-5.2-Codex's coding with stronger reasoning and runs 25% faster for Codex users.
Side by side
Specs and prices
| Fact | Claude Sonnet 4.6 | GPT-5.3-Codex |
|---|---|---|
| Maker | Anthropic | OpenAI |
| API model id | claude-sonnet-4-6 | gpt-5.3-codex |
| Released | February 17, 2026 | February 5, 2026 |
| Status | Previous generation | Previous generation |
| Context window | 1M tokens | 400K tokens |
| Max output | 128K tokens | 128K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $3 | $1.75 |
| Cache hit, per 1M | $0.30 | $0.175 |
| Cache write, per 1M | $3.75 (5-minute), $6 (1-hour) | $1.75 (same as input) |
| Output, per 1M tokens | $15 | $14 |
| Runs in | Claude Code, Cursor, OpenCode, and OpenRouter | Codex, Cursor, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker (Claude Sonnet 4.6: September 26, 2026; GPT-5.3-Codex: September 28, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Claude Sonnet 4.6: The full 1M context window is billed at standard rates on the API.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Claude Sonnet 4.6 | GPT-5.3-Codex |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $3.60 | $1.93 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.60 | $0.40 |
| Output-heavy generation, 30K input, 80K output | $1.29 | $1.17 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $396.00 | $211.75 |
| Where the session’s cost goes | ||
| Cache writes | $1.95 | $0.70 |
| Cache reads | $0.60 | $0.35 |
| Uncached input | $0.30 | $0.18 |
| Output | $0.75 | $0.70 |
- caching saves on the session with Claude Sonnet 4.6 (56%)
- $4.65
- caching saves on the session with GPT-5.3-Codex (62%)
- $3.15
Where GPT-5.3-Codex saves money, and where it doesn't
GPT-5.3-Codex charges $1.75 per million input tokens, $0.175 per million cached, and $14 per million output. Claude Sonnet 4.6 lists $3, $0.30, and $15 for the same three. Input and cache hits are 1.7x apart, but output is only 1.1x apart, $15 against $14.
Cache writes add to the input gap. OpenAI bills tokens written to the GPT-5.3-Codex cache as ordinary input at $1.75, while Anthropic charges $3.75 for a 5-minute write and $6 for a 1-hour write. On the example session, writes cost $1.95 on Sonnet 4.6 and $0.70 on GPT-5.3-Codex, and the session totals $3.60 against $1.93, a 1.9x gap. Over 110 sessions a month, that is $396.00 against $211.75 at API rates.
Output-heavy work is a different story. The generation workload, 30K in and 80K out, costs $1.29 on Sonnet 4.6 and $1.17 on GPT-5.3-Codex, so Sonnet 4.6 costs 10% more there, against 87% more on the session. Output is 36% of the GPT-5.3-Codex session and 21% of the Sonnet 4.6 session, so the more a task writes, the closer the two get.
How much context can each model take?
This is the sharpest difference. Sonnet 4.6 has a 1M context window and bills all of it at standard rates on the API. GPT-5.3-Codex has a 400K window, of which it accepts up to 272K as input, and both models write up to 128K tokens of output.
No request in the example session reaches 200K input tokens, so both models can run it. An agent that loads a large repository, long logs, or a big dependency tree into one request can pass 272K, and at that point GPT-5.3-Codex needs a smaller or compacted context while Sonnet 4.6 still has room.
Treat the limits as approximate when comparing. The two makers tokenize text differently, and Sonnet 4.6 uses Anthropic's older tokenizer, which counts fewer tokens for the same text than Claude models from 4.7 on. The same file won't be the same number of tokens on both.
Two February 2026 coding models, both now superseded
OpenAI released GPT-5.3-Codex on February 5, 2026, as a Codex-tuned coding model that combined GPT-5.2-Codex's coding with stronger reasoning and professional knowledge. OpenAI called it "The most capable agentic coding model to date" at the time, and credited it with collaborating better while the agent is working. GPT-5.4 later brought its coding capabilities into OpenAI's main model.
Anthropic released Sonnet 4.6 twelve days later, describing it as a full upgrade across coding, computer use, long-context reasoning, agent planning, knowledge work, and design. Anthropic now recommends moving to Claude Sonnet 5 and keeps Sonnet 4.6 on the API as a legacy model.
Access has narrowed on both sides. GPT-5.3-Codex has not been selectable in Codex with ChatGPT sign-in since May 26, 2026, though API-key use is unaffected, and it remains in Cursor, OpenRouter, OpenCode, and GitHub Copilot. Sonnet 4.6 left GitHub Copilot on September 1, 2026, apart from annual Copilot Pro and Pro+ subscribers, and it remains in Claude Code, Cursor, OpenRouter, and OpenCode.
Prompt caching
How each maker bills cached tokens
Anthropic
Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.
A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.
Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.
In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.
Source: Anthropic: Prompt caching
OpenAI
Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.
From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.
On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.
Source: OpenAI: Prompt caching
Your own numbers
See what Claude Sonnet 4.6 and GPT-5.3-Codex really cost you.
everyaitoken reads your Claude Code, Cursor, OpenCode, OpenRouter, and Codex history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is GPT-5.3-Codex cheaper than Claude Sonnet 4.6?
Yes, on every rate, though output is close at $14 against $15 per million. The example agentic session costs $1.93 on GPT-5.3-Codex and $3.60 on Sonnet 4.6.
What is GPT-5.3-Codex's context limit?
Its context window is 400K tokens, and it accepts up to 272K input tokens of that. Sonnet 4.6 accepts 1M.
Is GPT-5.3-Codex still available in Codex?
With an OpenAI API key, yes. It has not been selectable with ChatGPT sign-in since May 26, 2026.
Where can I see my own spend on each model?
EveryToken prices your local Claude Code, Codex, Cursor, OpenCode, and OpenRouter history at API rates, per model, and shows what caching saved or cost on each.
Sources
- Anthropic: Pricing
- Anthropic docs: Claude Sonnet 4.6
- Anthropic: Introducing Claude Sonnet 4.6
- Claude Code docs: Model configuration
- Cursor docs: Models and pricing
- OpenRouter: Claude Sonnet 4.6
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- OpenAI: API pricing
- OpenAI docs: GPT-5.3-Codex
- Codex docs: Changelog
- OpenAI: API changelog
- Codex docs: Models
- OpenRouter: GPT-5.3-Codex
- Anthropic: Prompt caching
- OpenAI: Prompt caching