Skip to content

Model comparison

GPT-5.4 vs GPT-5.3-Codex: does the 272K input cap matter?

GPT-5.3-Codex is cheaper per token but takes 272K input tokens at most. GPT-5.4 reaches 1.05M, at higher rates past 272K. Which one fits your codebase.

· Prices as of September 28, 2026

  • GPT-5.4

    OpenAI · Released March 5, 2026 · Previous generation

    The March 2026 flagship that brought GPT-5.3-Codex's coding into OpenAI's main model, now positioned as the more affordable option.

    GPT-5.4 facts and comparisons
  • GPT-5.3-Codex

    OpenAI · Released February 5, 2026 · Previous generation

    The February 2026 Codex-tuned coding model, which combined GPT-5.2-Codex's coding with stronger reasoning.

    GPT-5.3-Codex facts and comparisons

The short answer

GPT-5.3-Codex is the cheaper of the two, $1.93 against $2.50 for the example agentic coding session, because GPT-5.4 costs 43% more for input but only 7% more for output. The bigger difference is context: GPT-5.3-Codex accepts up to 272K input tokens, while GPT-5.4 has a 1.05M window with higher rates above 272K. OpenAI folded GPT-5.3-Codex's coding into GPT-5.4, and both have left Codex for ChatGPT sign-in.

Choose GPT-5.4 if

  • You need prompts beyond 272K input tokens, which GPT-5.4's 1.05M context window allows.
  • You want OpenAI's main model with GPT-5.3-Codex's coding built in, as OpenAI describes GPT-5.4.
  • You run long agent sessions that rely on compaction, a GPT-5.4 feature OpenAI highlights.

Choose GPT-5.3-Codex if

  • Your prompts stay under 272K input tokens and you want the lower rate: $1.75 input against $2.50.
  • You want the Codex-tuned model OpenAI called the most capable agentic coding model to date at its February 2026 launch.
  • Your work is output-heavy, where the example generation costs $1.17 against $1.28, a 9% gap.

Side by side

Specs and prices

FactGPT-5.4GPT-5.3-Codex
MakerOpenAIOpenAI
API model idgpt-5.4gpt-5.3-codex
ReleasedMarch 5, 2026February 5, 2026
StatusPrevious generationPrevious generation
Context window1.05M tokens400K tokens
Max output128K tokens128K tokens
Open weightsNoNo
Input, per 1M tokens$2.50$1.75
Cache hit, per 1M$0.25$0.175
Cache write, per 1M$2.50 (same as input)$1.75 (same as input)
Output, per 1M tokens$15$14
Runs inCodex, Cursor, OpenCode, OpenRouter, and GitHub CopilotCodex, Cursor, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. GPT-5.4: Prompts over 272K input tokens cost 2x for input and 1.5x for output, for the full session.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadGPT-5.4GPT-5.3-Codex
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$2.50$1.93
Large one-off review, 150K input with no cache hits, 10K output$0.53$0.40
Output-heavy generation, 30K input, 80K output$1.28$1.17
A month of sessions, 110 sessions: 5 a day, 22 working days$275.00$211.75
Where the session’s cost goes
Cache writes$1.00$0.70
Cache reads$0.50$0.35
Uncached input$0.25$0.18
Output$0.75$0.70
caching saves on the session with GPT-5.4 (64%)
$4.50
caching saves on the session with GPT-5.3-Codex (62%)
$3.15

Where GPT-5.3-Codex's 272K limit meets GPT-5.4's 1.05M window

GPT-5.3-Codex has a 400K context window, of which it accepts up to 272K input tokens. GPT-5.4 has a 1.05M window, about 2.6x as large, which OpenAI pitched for analyzing entire codebases.

The same number marks GPT-5.4's price step: a prompt above 272K input tokens moves the full session to 2x for input and 1.5x for output. Under 272K, both models work at their standard rates. Above it, only GPT-5.4 can take the request, and it charges a premium to do so.

The example session keeps every request under 272K, so it reflects the everyday case: $2.50 on GPT-5.4 and $1.93 on GPT-5.3-Codex. If your agents regularly load whole repositories into one prompt, the cap decides for you before price does.

Why the price gap depends on how much the model writes

Input is where the two part ways: $2.50 against $1.75 per million tokens, 43% more on GPT-5.4. Cache hits follow the same ratio, $0.25 against $0.175. Output is close, $15 against $14 per million, only 7% apart.

So the gap moves with the shape of the work. The uncached review, mostly input, costs $0.53 against $0.40. The output-heavy generation costs $1.28 against $1.17, 9% apart. The session sits in between at $2.50 against $1.93, a difference of $0.57.

Neither model charges extra to write the cache: both bill written tokens as ordinary input and charge 0.1x input for hits. In the session, cache writes cost $1.00 on GPT-5.4 and $0.70 on GPT-5.3-Codex. Over 110 sessions a month the totals are $275.00 and $211.75, a difference of $63.25.

What OpenAI says about GPT-5.3-Codex and GPT-5.4

OpenAI released GPT-5.3-Codex in February 2026 as a Codex-tuned coding model. It said the model brought GPT-5.2-Codex's coding together with stronger reasoning and professional knowledge, running 25% faster for Codex users, and collaborated better while the agent was working.

A month later, GPT-5.4 brought those coding capabilities into OpenAI's main model, according to OpenAI, along with the larger context window and compaction for longer agent runs. OpenAI now positions GPT-5.4 as its more affordable model for coding and professional work.

Both have left Codex for ChatGPT sign-in, GPT-5.3-Codex on May 26, 2026, and GPT-5.4 on August 31, 2026. Both still work in the API and in Codex with an API key. EveryToken reads local Codex history and prices every request at the published rates, so you can see which of the two your API-key sessions used and what each cost.

Prompt caching

How OpenAI bills cached tokens

OpenAI

Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.

From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.

On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.

Source: OpenAI: Prompt caching

Your own numbers

See what GPT-5.4 and GPT-5.3-Codex really cost you.

everyaitoken reads your Codex, Cursor, OpenCode, and OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

What is the context window of GPT-5.3-Codex?

GPT-5.3-Codex has a 400K context window and accepts up to 272K input tokens of it. It writes up to 128K tokens of output, the same as GPT-5.4.

Is GPT-5.3-Codex cheaper than GPT-5.4?

Yes. It charges $1.75 per million input tokens and $14 per million output tokens, against $2.50 and $15 for GPT-5.4. The example session costs $1.93 against $2.50.

Can I still use GPT-5.3-Codex in Codex?

Only with an API key. It has not been selectable in Codex with ChatGPT sign-in since May 26, 2026, and API-key use is unaffected.

Does GPT-5.4 charge more for long prompts?

Yes. Past 272K input tokens, GPT-5.4 bills input at 2x and output at 1.5x for the full session. GPT-5.3-Codex cannot accept prompts that large.

  • GPT-5.3-Codex vs GPT-5.2-Codex

    GPT-5.3-Codex and GPT-5.2-Codex cost the same to the cent. What differs is what OpenAI claims for each and where you can still run them in 2026.

  • GPT-5.5 vs GPT-5.4

    GPT-5.5 costs exactly twice GPT-5.4 on every rate, and both are leaving Codex sign-in. What the 2x gap means for API users and where Codex points instead.

  • Claude Sonnet 4.6 vs GPT-5.3-Codex

    Claude Sonnet 4.6 and GPT-5.3-Codex launched 12 days apart in February 2026. One takes 1M tokens of context, the other costs 46% less per session. Tradeoffs.

  • Claude Sonnet 4.6 vs GPT-5.4

    Claude Sonnet 4.6 and GPT-5.4 both charge $15 per million output tokens, so the cost gap lives in cache writes. What that means for sessions and upgrades.

  • GPT-5.3-Codex vs Gemini 3.1 Pro Preview

    GPT-5.3-Codex and Gemini 3.1 Pro Preview cost within 4% of each other on a cached coding session. Context size, output limits, and access set them apart.

  • GPT-5.6 Sol vs GPT-5.5

    GPT-5.6 Sol cuts GPT-5.5's rates but adds a 1.25x charge for cache writes. Why the session gap is 16%, not 20%, and what changes on October 14.