Skip to content

Model comparison

GLM-5.3-Flash vs GPT-6 Luna: cache hits decide the price

GPT-6 Luna and GLM-5.3-Flash share a $0.50 output rate, yet a cached coding session costs $0.11 on Luna and $0.16 on GLM-5.3-Flash. Cache hits explain why.

· Prices as of September 28, 2026

  • GLM-5.3-Flash

    Z.ai · Released August 26, 2026

    Z.ai's low-cost, natively multimodal model, which Z.ai says outperforms GLM-5.2 at a tenth of the price.

    GLM-5.3-Flash facts and comparisons
  • GPT-6 Luna

    OpenAI · Released September 22, 2026

    The cheapest GPT-6 model, pitched for focused, high-volume, repeatable work, including narrower coding tasks.

    GPT-6 Luna facts and comparisons

The short answer

GPT-6 Luna is the cheaper of the two at list prices: the example agentic coding session costs $0.11 against $0.16 on GLM-5.3-Flash, and most of the $0.05 gap is cache reads, where Luna charges $0.01 per million against $0.03. Both charge $0.50 per million output tokens, so output-heavy work costs the same $0.04 on each. Pick GPT-6 Luna for Codex or GitHub Copilot, and GLM-5.3-Flash for open MIT-licensed weights, visual coding, or work through OpenRouter and OpenCode.

Choose GLM-5.3-Flash if

  • You want open weights under the MIT license that you can host yourself.
  • Your agent checks its own UI work, which Z.ai supports with visual coding that looks at interfaces and rendered results.
  • Most of your spend is output, where GLM-5.3-Flash's $0.50 per million matches GPT-6 Luna's.
  • You already route coding traffic through OpenRouter, where GLM-5.3-Flash was the most-used programming model over the week before September 28, 2026, summed across nine languages.

Choose GPT-6 Luna if

  • Your agent rereads a large cached context, and Luna's $0.01 hits cost a third of GLM-5.3-Flash's $0.03.
  • You code in Codex, whose docs point to GPT-6 Luna for focused, repeatable tasks.
  • You use GitHub Copilot, or the Codex app on a Free or Go plan, which include GPT-6 Luna.
  • You send lots of fresh input, where Luna's $0.10 per million undercuts $0.15.

Side by side

Specs and prices

FactGLM-5.3-FlashGPT-6 Luna
MakerZ.aiOpenAI
API model idglm-5.3-flashgpt-6-luna
ReleasedAugust 26, 2026September 22, 2026
StatusCurrentCurrent
Context window1M tokens1.05M tokens
Max output128K tokens128K tokens
Open weightsYesNo
Input, per 1M tokens$0.15$0.10
Cache hit, per 1M$0.03$0.01
Cache write, per 1M$0.15 (same as input)$0.125
Output, per 1M tokens$0.50$0.50
Runs inOpenCode and OpenRouterCodex, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. GPT-6 Luna: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadGLM-5.3-FlashGPT-6 Luna
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$0.16$0.11
Large one-off review, 150K input with no cache hits, 10K output$0.03$0.02
Output-heavy generation, 30K input, 80K output$0.04$0.04
A month of sessions, 110 sessions: 5 a day, 22 working days$17.60$11.55
Where the session’s cost goes
Cache writes$0.06$0.05
Cache reads$0.06$0.02
Uncached input$0.02$0.01
Output$0.03$0.03
caching saves on the session with GLM-5.3-Flash (60%)
$0.24
caching saves on the session with GPT-6 Luna (61%)
$0.17

Same output price, different cache discounts

GPT-6 Luna and GLM-5.3-Flash both charge $0.50 per million output tokens, so the output-heavy generation costs $0.04 on each. Input is 1.5x apart, $0.10 on Luna and $0.15 on GLM-5.3-Flash, and the large one-off review follows it at $0.02 against $0.03.

The session gap comes mostly from the cache. OpenAI bills a hit at 10% of input, $0.01 on Luna. Z.ai bills $0.03, 20% of GLM-5.3-Flash's input price. The example session reads 2M tokens from the cache, which costs $0.02 on Luna and $0.06 on GLM-5.3-Flash, and that line alone is $0.04 of the $0.05 gap.

Cache writes almost cancel out. OpenAI charges 1.25x input to write the cache from GPT-5.6 on, $0.125 per million on Luna, and Z.ai lists no write fee, so GLM-5.3-Flash's written tokens cost its $0.15 input rate. Even with OpenAI's premium, Luna's writes come in lower: $0.05 against $0.06 for the session's 400K written tokens.

What $0.05 a session adds up to

At 110 sessions a month, five each working day, the example comes to $11.55 on GPT-6 Luna and $17.60 on GLM-5.3-Flash. That is 52% more on GLM-5.3-Flash, but in dollars it is $6.05 a month.

Caching does similar work on both. It saves 61% of the uncached session cost on Luna and 60% on GLM-5.3-Flash: Z.ai's shallower hit discount is offset by writes that carry no premium, and OpenAI's deeper discount is offset by the 1.25x write price.

Rates on OpenRouter can differ from both lists. It sends a request for an open-weight model like GLM-5.3-Flash to one of several providers, each with its own price. The tables here use Z.ai's and OpenAI's own API prices.

How OpenAI and Z.ai position these models

In OpenAI's words, GPT-6 Luna is "our most efficient model for focused, high-volume tasks." It is the model Codex's documentation points to for narrow jobs that repeat. Its reasoning effort starts at medium, and Codex lets it go up to max. Codex, OpenAI's own agent, offers it, though not in Codex cloud, and Free and Go plans get it in the Codex app.

Z.ai describes GLM-5.3-Flash as a low-cost, natively multimodal model that outperforms GLM-5.2 at one-tenth the price. It highlights visual coding, where the model looks at interfaces and rendered results to test and improve its work, and hybrid sparse and linear attention that shrinks attention compute and KV cache size. The weights are MIT-licensed, so self-hosting is an option this page does not price.

Both have large windows: 1.05M tokens of context on Luna and 1M on GLM-5.3-Flash, with 128K of output on each. Luna has a long-context rule worth knowing if your agent loads whole repositories: past 272K input tokens, OpenAI bills the entire request at 2x for input and cache and 1.5x for output.

EveryToken prices GPT-6 Luna at OpenAI's rates in Codex and OpenCode, and prices GLM-5.3-Flash when you use it through OpenRouter.

Prompt caching

How each maker bills cached tokens

Z.ai

Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.

A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.

Source: Z.ai docs: Context caching

OpenAI

Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.

From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.

On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.

Source: OpenAI: Prompt caching

Your own numbers

See what GLM-5.3-Flash and GPT-6 Luna really cost you.

everyaitoken reads your OpenRouter, Codex, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is GPT-6 Luna cheaper than GLM-5.3-Flash?

At list prices, yes, on input and cache hits, while output costs $0.50 per million on both. The example cached session comes to $0.11 on GPT-6 Luna and $0.16 on GLM-5.3-Flash.

Why does GLM-5.3-Flash cost more on a cached session?

Z.ai prices a cache hit at 20% of input, $0.03 per million, while OpenAI prices it at 10%, $0.01 on Luna. Reading the session's 2M cached tokens costs $0.06 against $0.02.

Where can I use GLM-5.3-Flash?

Through OpenRouter and OpenCode, or on Z.ai's API. It is not in Cursor or GitHub Copilot, and its MIT-licensed weights also allow self-hosting.

What happens above 272K input tokens on GPT-6 Luna?

OpenAI switches the whole request to 2x for input and cache and 1.5x for output. Each request in the example session stays below 200K, so the tables use Luna's standard rates.

  • GLM-5.3 vs GLM-5.3-Flash

    GLM-5.3-Flash costs about a ninth of GLM-5.3 at Z.ai's rates: $0.16 against $1.44 for an agentic coding session. Same limits, same caching, different jobs.

  • GPT-6 Astra vs GPT-6 Luna

    GPT-6 Astra and GPT-6 Luna sit at opposite ends of OpenAI's lineup, 100x apart per token. What each is for, and where GPT-6 Sol fits between them.

  • GPT-6 Sol vs GPT-6 Luna

    GPT-6 Luna costs a twentieth of GPT-6 Sol per token. What OpenAI and the Codex docs say each tier is for, and what the gap means for coding sessions.

  • GPT-6 Luna vs GPT-5.6 Luna

    GPT-6 Luna halves the input price of GPT-5.6 Luna, which already had an 80% cut, and trims output further. What the move saves, and what Codex suggests.

  • Claude Haiku 4.5 vs GPT-6 Luna

    GPT-6 Luna lists a tenth of Claude Haiku 4.5's rates and a 1.05M context window against 200K. What each costs for sub-agents and cached coding sessions.

  • GLM-5.3-Flash vs Claude Sonnet 5

    GLM-5.3-Flash costs $0.16 for a cached coding session that costs $2.40 on Claude Sonnet 5. What drives a 15x gap, and what price alone can't tell you.