Model comparison
GPT-5.4 vs GPT-5.3-Codex: does the 272K input cap matter?
GPT-5.3-Codex is cheaper per token but takes 272K input tokens at most. GPT-5.4 reaches 1.05M, at higher rates past 272K. Which one fits your codebase.
· Prices as of September 28, 2026
GPT-5.4
OpenAI · Released March 5, 2026 · Previous generation
The March 2026 flagship that brought GPT-5.3-Codex's coding into OpenAI's main model, now positioned as the more affordable option.
GPT-5.4 facts and comparisonsGPT-5.3-Codex
OpenAI · Released February 5, 2026 · Previous generation
The February 2026 Codex-tuned coding model, which combined GPT-5.2-Codex's coding with stronger reasoning.
GPT-5.3-Codex facts and comparisons
The short answer
GPT-5.3-Codex is the cheaper of the two, $1.93 against $2.50 for the example agentic coding session, because GPT-5.4 costs 43% more for input but only 7% more for output. The bigger difference is context: GPT-5.3-Codex accepts up to 272K input tokens, while GPT-5.4 has a 1.05M window with higher rates above 272K. OpenAI folded GPT-5.3-Codex's coding into GPT-5.4, and both have left Codex for ChatGPT sign-in.
Choose GPT-5.4 if
- You need prompts beyond 272K input tokens, which GPT-5.4's 1.05M context window allows.
- You want OpenAI's main model with GPT-5.3-Codex's coding built in, as OpenAI describes GPT-5.4.
- You run long agent sessions that rely on compaction, a GPT-5.4 feature OpenAI highlights.
Choose GPT-5.3-Codex if
- Your prompts stay under 272K input tokens and you want the lower rate: $1.75 input against $2.50.
- You want the Codex-tuned model OpenAI called the most capable agentic coding model to date at its February 2026 launch.
- Your work is output-heavy, where the example generation costs $1.17 against $1.28, a 9% gap.
Side by side
Specs and prices
| Fact | GPT-5.4 | GPT-5.3-Codex |
|---|---|---|
| Maker | OpenAI | OpenAI |
| API model id | gpt-5.4 | gpt-5.3-codex |
| Released | March 5, 2026 | February 5, 2026 |
| Status | Previous generation | Previous generation |
| Context window | 1.05M tokens | 400K tokens |
| Max output | 128K tokens | 128K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $2.50 | $1.75 |
| Cache hit, per 1M | $0.25 | $0.175 |
| Cache write, per 1M | $2.50 (same as input) | $1.75 (same as input) |
| Output, per 1M tokens | $15 | $14 |
| Runs in | Codex, Cursor, OpenCode, OpenRouter, and GitHub Copilot | Codex, Cursor, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. GPT-5.4: Prompts over 272K input tokens cost 2x for input and 1.5x for output, for the full session.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | GPT-5.4 | GPT-5.3-Codex |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $2.50 | $1.93 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.53 | $0.40 |
| Output-heavy generation, 30K input, 80K output | $1.28 | $1.17 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $275.00 | $211.75 |
| Where the session’s cost goes | ||
| Cache writes | $1.00 | $0.70 |
| Cache reads | $0.50 | $0.35 |
| Uncached input | $0.25 | $0.18 |
| Output | $0.75 | $0.70 |
- caching saves on the session with GPT-5.4 (64%)
- $4.50
- caching saves on the session with GPT-5.3-Codex (62%)
- $3.15
Where GPT-5.3-Codex's 272K limit meets GPT-5.4's 1.05M window
GPT-5.3-Codex has a 400K context window, of which it accepts up to 272K input tokens. GPT-5.4 has a 1.05M window, about 2.6x as large, which OpenAI pitched for analyzing entire codebases.
The same number marks GPT-5.4's price step: a prompt above 272K input tokens moves the full session to 2x for input and 1.5x for output. Under 272K, both models work at their standard rates. Above it, only GPT-5.4 can take the request, and it charges a premium to do so.
The example session keeps every request under 272K, so it reflects the everyday case: $2.50 on GPT-5.4 and $1.93 on GPT-5.3-Codex. If your agents regularly load whole repositories into one prompt, the cap decides for you before price does.
Why the price gap depends on how much the model writes
Input is where the two part ways: $2.50 against $1.75 per million tokens, 43% more on GPT-5.4. Cache hits follow the same ratio, $0.25 against $0.175. Output is close, $15 against $14 per million, only 7% apart.
So the gap moves with the shape of the work. The uncached review, mostly input, costs $0.53 against $0.40. The output-heavy generation costs $1.28 against $1.17, 9% apart. The session sits in between at $2.50 against $1.93, a difference of $0.57.
Neither model charges extra to write the cache: both bill written tokens as ordinary input and charge 0.1x input for hits. In the session, cache writes cost $1.00 on GPT-5.4 and $0.70 on GPT-5.3-Codex. Over 110 sessions a month the totals are $275.00 and $211.75, a difference of $63.25.
What OpenAI says about GPT-5.3-Codex and GPT-5.4
OpenAI released GPT-5.3-Codex in February 2026 as a Codex-tuned coding model. It said the model brought GPT-5.2-Codex's coding together with stronger reasoning and professional knowledge, running 25% faster for Codex users, and collaborated better while the agent was working.
A month later, GPT-5.4 brought those coding capabilities into OpenAI's main model, according to OpenAI, along with the larger context window and compaction for longer agent runs. OpenAI now positions GPT-5.4 as its more affordable model for coding and professional work.
Both have left Codex for ChatGPT sign-in, GPT-5.3-Codex on May 26, 2026, and GPT-5.4 on August 31, 2026. Both still work in the API and in Codex with an API key. EveryToken reads local Codex history and prices every request at the published rates, so you can see which of the two your API-key sessions used and what each cost.
Prompt caching
How OpenAI bills cached tokens
OpenAI
Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.
From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.
On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.
Source: OpenAI: Prompt caching
Your own numbers
See what GPT-5.4 and GPT-5.3-Codex really cost you.
everyaitoken reads your Codex, Cursor, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
What is the context window of GPT-5.3-Codex?
GPT-5.3-Codex has a 400K context window and accepts up to 272K input tokens of it. It writes up to 128K tokens of output, the same as GPT-5.4.
Is GPT-5.3-Codex cheaper than GPT-5.4?
Yes. It charges $1.75 per million input tokens and $14 per million output tokens, against $2.50 and $15 for GPT-5.4. The example session costs $1.93 against $2.50.
Can I still use GPT-5.3-Codex in Codex?
Only with an API key. It has not been selectable in Codex with ChatGPT sign-in since May 26, 2026, and API-key use is unaffected.
Does GPT-5.4 charge more for long prompts?
Yes. Past 272K input tokens, GPT-5.4 bills input at 2x and output at 1.5x for the full session. GPT-5.3-Codex cannot accept prompts that large.