Skip to content

Model comparison

GLM-5.3 vs GPT-6 Sol: open weights or Codex's coding pick

GLM-5.3 costs 31% less than GPT-6 Sol on an agentic coding session, $1.44 against $2.10. How OpenAI's write fee and 272K rule compare with Z.ai's rates.

· Prices as of September 28, 2026

  • GLM-5.3

    Z.ai · Released August 14, 2026

    Z.ai's flagship for complex software engineering and long-horizon agent work, with open weights.

    GLM-5.3 facts and comparisons
  • GPT-6 Sol

    OpenAI · Released September 22, 2026

    The mid-priced GPT-6 model, which OpenAI pitches for complex coding and agent workflows and which the Codex docs recommend for complex coding.

    GPT-6 Sol facts and comparisons

The short answer

GLM-5.3 costs $1.44 on the example agentic coding session against $2.10 on GPT-6 Sol, 31% less, and the gap reaches 55% on output-heavy work. Pick GPT-6 Sol if you code in Codex, whose docs recommend it for complex coding; pick GLM-5.3 if you want lower rates and open weights, through OpenRouter or OpenCode.

Choose GLM-5.3 if

  • You want the lower rates: $1.40 input and $4.40 output per million tokens, against $2 and $10 on GPT-6 Sol.
  • Your agent writes a lot to the cache, which Z.ai bills at the input rate while OpenAI charges 1.25x input for a write.
  • You want the option to self-host, which GLM-5.3's open weights allow under Z.ai's own license.
  • Your client speaks an OpenAI or Anthropic API format: Z.ai serves GLM-5.3 on Chat Completions, Responses, and Anthropic Messages endpoints.

Choose GPT-6 Sol if

  • You are upgrading within Codex, whose docs suggest GPT-6 Sol as the move from GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.4.
  • You want to steer caching yourself, with up to four explicit breakpoints and a prefix that stays reusable for at least 30 minutes after its last use.
  • You choose models in GitHub Copilot, which offers GPT-6 Sol and does not offer GLM-5.3.
  • Your sessions reread their context heavily, where a GPT-6 Sol cache hit costs $0.20 per million against $0.26 on GLM-5.3, though its writes and output still cost more.

Side by side

Specs and prices

FactGLM-5.3GPT-6 Sol
MakerZ.aiOpenAI
API model idglm-5.3gpt-6-sol
ReleasedAugust 14, 2026September 22, 2026
StatusCurrentCurrent
Context window1M tokens1.05M tokens
Max output128K tokens128K tokens
Open weightsYesNo
Input, per 1M tokens$1.40$2
Cache hit, per 1M$0.26$0.20
Cache write, per 1M$1.40 (same as input)$2.50
Output, per 1M tokens$4.40$10
Runs inOpenCode and OpenRouterCodex, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. GPT-6 Sol: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadGLM-5.3GPT-6 Sol
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$1.44$2.10
Large one-off review, 150K input with no cache hits, 10K output$0.25$0.40
Output-heavy generation, 30K input, 80K output$0.39$0.86
A month of sessions, 110 sessions: 5 a day, 22 working days$158.40$231.00
Where the session’s cost goes
Cache writes$0.56$1.00
Cache reads$0.52$0.40
Uncached input$0.14$0.20
Output$0.22$0.50
caching saves on the session with GLM-5.3 (61%)
$2.28
caching saves on the session with GPT-6 Sol (62%)
$3.40

How much more does GPT-6 Sol cost than GLM-5.3?

GPT-6 Sol lists $2 per million input tokens and $10 per million output tokens. GLM-5.3 lists $1.40 and $4.40. On uncached work GPT-6 Sol costs 1.6x as much for the large one-off review, $0.40 against $0.25, and 2.2x as much for the output-heavy generation, $0.86 against $0.39.

The agentic coding session narrows that to 1.5x: $2.10 on GPT-6 Sol and $1.44 on GLM-5.3, a $0.66 difference. Two lines drive it. OpenAI charges 1.25x input to write the cache, $2.50 per million, so 400K written tokens cost $1.00 against $0.56 on GLM-5.3, which bills writes as ordinary input. Output adds $0.28 more on GPT-6 Sol.

One line goes the other way. The session reads 2M tokens from the cache, and GPT-6 Sol charges $0.20 per million for a hit while GLM-5.3 charges $0.26, so reads cost $0.40 against $0.52. Over 110 sessions a month the totals are $231.00 and $158.40, a $72.60 difference at each maker's published rates.

Caching rules at OpenAI and Z.ai

OpenAI turns prompt caching on by default. For GPT-6 Sol that means up to four explicit breakpoints, a cached prefix that stays reusable for at least 30 minutes after its last use, and caching from 1,024 input tokens. On GPT-6 Sol, the cost of that control is a write charge of 1.25x the uncached input price.

GLM-5.3's cache is automatic, with no configuration and no write fee. Its discount is shallower, though: a hit costs 18.6% of the GLM-5.3 input price, where a GPT-6 Sol hit costs 10% of its input. In percentage terms the two effects nearly cancel. Caching saves 62% on GPT-6 Sol and 61% on GLM-5.3 against sending the same tokens uncached, or $3.40 and $2.28.

The cache section on our homepage shows how the same session's reads and writes add up. On your own work, the ratio of cache writes to cache reads decides which maker's rules cost less.

Context limits, the 272K line, and reasoning effort

GPT-6 Sol has a 1.05M context window and accepts up to 922K input tokens of it. GLM-5.3 has 1M. Both write up to 128K tokens of output. Once a GPT-6 Sol request passes 272K input tokens, all of it is billed at 2x for input and cache and 1.5x for output, while the GLM-5.3 rates used here carry no long-context tier.

The example session keeps each request under 200K tokens, so neither GLM-5.3 nor GPT-6 Sol changes rates on it. An agent that loads a large repository or a long log into one request can cross 272K on GPT-6 Sol, and then every token in that request is billed at the higher rates.

Reasoning defaults move cost too. GPT-6 Sol runs at medium effort by default, in the API and in Codex. GLM-5.3 always reasons, at low, high, or max. More reasoning means more output tokens, and output is where the list prices differ most, so compare the two at the effort levels you would actually use.

Where GLM-5.3 and GPT-6 Sol run

Codex is OpenAI's own coding agent, built around OpenAI's models, and its docs recommend GPT-6 Sol for complex coding. Beyond Codex, GPT-6 Sol is offered in OpenCode, through OpenRouter, and in GitHub Copilot. On its model page, OpenAI says GPT-6 Sol is "built to power complex coding and agentic workflows."

GLM-5.3 is offered through OpenRouter and in OpenCode. Z.ai calls it its latest flagship for complex software engineering and agent capabilities, and publishes open weights under its own GLM-5.3 license. The GLM-5.3 prices in this comparison are Z.ai's own, and OpenRouter providers hosting the weights can charge differently.

EveryToken prices GPT-6 Sol at OpenAI's rates from Codex and OpenCode history, and prices GLM-5.3 from OpenRouter's catalog when you use it through OpenRouter. The GPT-6 Sol figures are API-equivalent estimates at published rates, not what a ChatGPT plan charges for it.

Prompt caching

How each maker bills cached tokens

Z.ai

Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.

A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.

Source: Z.ai docs: Context caching

OpenAI

Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.

From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.

On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.

Source: OpenAI: Prompt caching

Your own numbers

See what GLM-5.3 and GPT-6 Sol really cost you.

everyaitoken reads your OpenRouter, Codex, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is GLM-5.3 cheaper than GPT-6 Sol?

Yes. On the example agentic session it costs $1.44 against $2.10, and on output-heavy generation $0.39 against $0.86. GPT-6 Sol charges less for a cache hit, $0.20 against $0.26 per million, but its write charge outweighs that on the session.

Can I use GLM-5.3 in Codex?

Codex is built around OpenAI's models and defaults to them, and its docs recommend GPT-6 Sol for complex coding. GLM-5.3 is available through OpenRouter and in OpenCode, and Z.ai's API accepts OpenAI Responses requests; connecting Codex to another provider takes custom configuration that this page does not cover.

Does GPT-6 Sol charge more for long prompts?

Yes. A GPT-6 Sol prompt above 272K input tokens moves the entire request to 2x for input and cache and 1.5x for output. The model takes up to 922K input tokens of its 1.05M window, and GLM-5.3's rates here have no comparable step.

Which model has the larger context window?

GPT-6 Sol, at 1.05M tokens against 1M on GLM-5.3, a 5% difference. Both write up to 128K tokens of output. For cost, GPT-6 Sol's 272K pricing line matters more than the size of the window.

  • GLM-5.3 vs GLM-5.3-Flash

    GLM-5.3-Flash costs about a ninth of GLM-5.3 at Z.ai's rates: $0.16 against $1.44 for an agentic coding session. Same limits, same caching, different jobs.

  • GPT-6 Astra vs GPT-6 Sol

    GPT-6 Astra costs 5x as much as GPT-6 Sol on every token, cached or not. What OpenAI built each for, which one Codex picks, and what a session costs.

  • GPT-6 Sol vs GPT-5.6 Sol

    GPT-6 Sol costs half as much as GPT-5.6 Sol on every rate, and Codex suggests the move. What changes, what the promo pricing means, and what stays put.

  • GPT-6 Sol vs GPT-5.6 Terra

    GPT-6 Sol matches GPT-5.6 Terra's input and cache prices and charges less for output. What that means for coding sessions and the move Codex suggests.

  • GPT-6 Sol vs GPT-6 Luna

    GPT-6 Luna costs a twentieth of GPT-6 Sol per token. What OpenAI and the Codex docs say each tier is for, and what the gap means for coding sessions.

  • Claude Fable 5.1 vs GPT-6 Sol

    Claude Fable 5.1 costs 5x what GPT-6 Sol does per token, and caching leaves that ratio intact. How Anthropic's top model compares with OpenAI's middle one.