Model comparison
Claude Sonnet 4.6 vs GPT-5.4: API cost for coding agents
Claude Sonnet 4.6 and GPT-5.4 both charge $15 per million output tokens, so the cost gap lives in cache writes. What that means for sessions and upgrades.
· Prices as of September 28, 2026
Claude Sonnet 4.6
Anthropic · Released February 17, 2026 · Previous generation
The previous Sonnet, pitched at launch as near-Opus capability at the Sonnet price. Anthropic now recommends moving to Claude Sonnet 5.
Claude Sonnet 4.6 facts and comparisonsGPT-5.4
OpenAI · Released March 5, 2026 · Previous generation
The March 2026 flagship that brought GPT-5.3-Codex's coding into OpenAI's main model, now positioned as the more affordable option.
GPT-5.4 facts and comparisons
The short answer
GPT-5.4 costs less on the example agentic coding session, $2.50 against $3.60 on Claude Sonnet 4.6, mostly because OpenAI bills cache writes as ordinary input. Output costs $15 per million on both, so output-heavy work is nearly a tie at $1.28 and $1.29. Both are previous-generation models: Anthropic recommends Claude Sonnet 5, and GPT-5.4 left Codex's ChatGPT sign-in on August 31, 2026.
Choose Claude Sonnet 4.6 if
- You work in Claude Code and have prompts validated on Sonnet 4.6, which stays available on the API.
- Your prompts run past 272K input tokens, where GPT-5.4 charges 2x input and 1.5x output for the full session and Sonnet 4.6 bills its full 1M window at standard rates on the API.
- Your sessions pause long enough between turns to want a 1-hour cache, which Anthropic offers at 2x input.
Choose GPT-5.4 if
- You run cache-heavy agentic sessions and want to avoid a write premium: GPT-5.4 bills written tokens at its $2.50 input rate.
- You use Codex with an API key, where GPT-5.4 remains usable.
- You want the model OpenAI now positions as its more affordable option for coding and professional work.
Side by side
Specs and prices
| Fact | Claude Sonnet 4.6 | GPT-5.4 |
|---|---|---|
| Maker | Anthropic | OpenAI |
| API model id | claude-sonnet-4-6 | gpt-5.4 |
| Released | February 17, 2026 | March 5, 2026 |
| Status | Previous generation | Previous generation |
| Context window | 1M tokens | 1.05M tokens |
| Max output | 128K tokens | 128K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $3 | $2.50 |
| Cache hit, per 1M | $0.30 | $0.25 |
| Cache write, per 1M | $3.75 (5-minute), $6 (1-hour) | $2.50 (same as input) |
| Output, per 1M tokens | $15 | $15 |
| Runs in | Claude Code, Cursor, OpenCode, and OpenRouter | Codex, Cursor, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker (Claude Sonnet 4.6: September 26, 2026; GPT-5.4: September 28, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Claude Sonnet 4.6: The full 1M context window is billed at standard rates on the API. GPT-5.4: Prompts over 272K input tokens cost 2x for input and 1.5x for output, for the full session.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Claude Sonnet 4.6 | GPT-5.4 |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $3.60 | $2.50 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.60 | $0.53 |
| Output-heavy generation, 30K input, 80K output | $1.29 | $1.28 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $396.00 | $275.00 |
| Where the session’s cost goes | ||
| Cache writes | $1.95 | $1.00 |
| Cache reads | $0.60 | $0.50 |
| Uncached input | $0.30 | $0.25 |
| Output | $0.75 | $0.75 |
- caching saves on the session with Claude Sonnet 4.6 (56%)
- $4.65
- caching saves on the session with GPT-5.4 (64%)
- $4.50
Why GPT-5.4 costs 31% less on an agentic session
The headline rates are close. Claude Sonnet 4.6 charges $3 per million input tokens and GPT-5.4 charges $2.50. Cache hits are $0.30 and $0.25. Output is $15 per million on both. On the output-heavy generation, which is mostly output, that makes the two nearly identical: $1.29 on Sonnet 4.6 and $1.28 on GPT-5.4.
Cache writes are where they part. Anthropic bills a 5-minute write at $3.75 and a 1-hour write at $6. OpenAI adds no charge for writing the cache on GPT-5.4 and bills those tokens as ordinary input at $2.50. The example session writes 400K tokens, so writes cost $1.95 on Sonnet 4.6, 54% of its session, against $1.00 on GPT-5.4. That $0.95 is most of the $1.10 gap between $3.60 and $2.50.
The large one-off review, with no caching at all, shows the plain input gap: $0.60 against $0.53. Over 110 sessions a month the projection is $396.00 on Sonnet 4.6 against $275.00 on GPT-5.4 at API rates.
Long prompts, context windows, and caching rules
Both accept roughly a million tokens: 1M on Sonnet 4.6 and 1.05M on GPT-5.4, each with 128K of output. Pricing past 272K input tokens differs. GPT-5.4 charges 2x for input and 1.5x for output once a prompt crosses that line, for the full session, while Sonnet 4.6 bills its full 1M window at standard rates on the API. OpenAI pitches GPT-5.4's 1M context window for analyzing entire codebases, and supports compaction for longer agent runs.
Caching works differently on the two. GPT-5.4 caches automatically, with no explicit breakpoints, and a hit costs 0.1x input. Claude caches a prompt prefix up to a breakpoint that you place or that one top-level cache_control field moves for you, with 5-minute and 1-hour lifetimes. In Claude Code, the main conversation uses the 1-hour cache on a subscription and the 5-minute cache with an API key.
Sonnet 4.6 uses Anthropic's older tokenizer, and the two makers tokenize text differently, so the same prompt counts as a different number of tokens on each. The fixed workloads here hold token counts equal, which real prompts won't.
Both models now have successors
Anthropic released Sonnet 4.6 in February 2026, calling it "a full upgrade of the model's skills across coding, computer use, long-context reasoning, agent planning, knowledge work, and design," and pitched it as approaching Opus-level intelligence at a more practical price. It is now a legacy model that stays available on the API, and Anthropic recommends moving to Claude Sonnet 5. GitHub Copilot retired Sonnet 4.6 on September 1, 2026, except for annual Copilot Pro and Pro+ subscribers.
GPT-5.4 arrived in March 2026 and brought GPT-5.3-Codex's coding capabilities into OpenAI's main model. OpenAI now describes it as "A more affordable model for coding and professional work." It was retired from Codex for ChatGPT sign-in on August 31, 2026, and remains usable in Codex with an API key. For Codex users, the suggested next model is GPT-6 Sol.
Prompt caching
How each maker bills cached tokens
Anthropic
Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.
A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.
Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.
In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.
Source: Anthropic: Prompt caching
OpenAI
Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.
From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.
On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.
Source: OpenAI: Prompt caching
Your own numbers
See what Claude Sonnet 4.6 and GPT-5.4 really cost you.
everyaitoken reads your Claude Code, Cursor, OpenCode, OpenRouter, and Codex history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is GPT-5.4 cheaper than Claude Sonnet 4.6?
For input and cache writes, yes. Output costs $15 per million on both. The example agentic session is $2.50 on GPT-5.4 and $3.60 on Sonnet 4.6, mostly because GPT-5.4 has no cache-write premium.
Can I still use GPT-5.4 in Codex?
Yes, with an API key. Signing in with ChatGPT no longer offers it in Codex, as of August 31, 2026.
Is Claude Sonnet 4.6 still available?
On the Anthropic API, yes, as a legacy model, and it is still listed in Claude Code, Cursor, OpenRouter, and OpenCode. Anthropic points Sonnet 4.6 users to Claude Sonnet 5.
How can I compare these costs on my own work?
EveryToken, a $9 one-time Mac app, reads local Claude Code and Codex history and prices each request at API rates, with cache writes and hits broken out, so the effect of the write premium shows in your own numbers.
Sources
- Anthropic: Pricing
- Anthropic docs: Claude Sonnet 4.6
- Anthropic: Introducing Claude Sonnet 4.6
- Claude Code docs: Model configuration
- Cursor docs: Models and pricing
- OpenRouter: Claude Sonnet 4.6
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- OpenAI: API pricing
- OpenAI docs: GPT-5.4
- OpenAI: Using GPT-5.4
- OpenAI: API changelog
- Codex docs: Models
- OpenRouter: GPT-5.4
- Anthropic: Prompt caching
- OpenAI: Prompt caching