Model comparison
GLM-5.3-Flash vs GPT-6 Luna: cache hits decide the price
GPT-6 Luna and GLM-5.3-Flash share a $0.50 output rate, yet a cached coding session costs $0.11 on Luna and $0.16 on GLM-5.3-Flash. Cache hits explain why.
· Prices as of September 28, 2026
GLM-5.3-Flash
Z.ai · Released August 26, 2026
Z.ai's low-cost, natively multimodal model, which Z.ai says outperforms GLM-5.2 at a tenth of the price.
GLM-5.3-Flash facts and comparisonsGPT-6 Luna
OpenAI · Released September 22, 2026
The cheapest GPT-6 model, pitched for focused, high-volume, repeatable work, including narrower coding tasks.
GPT-6 Luna facts and comparisons
The short answer
GPT-6 Luna is the cheaper of the two at list prices: the example agentic coding session costs $0.11 against $0.16 on GLM-5.3-Flash, and most of the $0.05 gap is cache reads, where Luna charges $0.01 per million against $0.03. Both charge $0.50 per million output tokens, so output-heavy work costs the same $0.04 on each. Pick GPT-6 Luna for Codex or GitHub Copilot, and GLM-5.3-Flash for open MIT-licensed weights, visual coding, or work through OpenRouter and OpenCode.
Choose GLM-5.3-Flash if
- You want open weights under the MIT license that you can host yourself.
- Your agent checks its own UI work, which Z.ai supports with visual coding that looks at interfaces and rendered results.
- Most of your spend is output, where GLM-5.3-Flash's $0.50 per million matches GPT-6 Luna's.
- You already route coding traffic through OpenRouter, where GLM-5.3-Flash was the most-used programming model over the week before September 28, 2026, summed across nine languages.
Choose GPT-6 Luna if
- Your agent rereads a large cached context, and Luna's $0.01 hits cost a third of GLM-5.3-Flash's $0.03.
- You code in Codex, whose docs point to GPT-6 Luna for focused, repeatable tasks.
- You use GitHub Copilot, or the Codex app on a Free or Go plan, which include GPT-6 Luna.
- You send lots of fresh input, where Luna's $0.10 per million undercuts $0.15.
Side by side
Specs and prices
| Fact | GLM-5.3-Flash | GPT-6 Luna |
|---|---|---|
| Maker | Z.ai | OpenAI |
| API model id | glm-5.3-flash | gpt-6-luna |
| Released | August 26, 2026 | September 22, 2026 |
| Status | Current | Current |
| Context window | 1M tokens | 1.05M tokens |
| Max output | 128K tokens | 128K tokens |
| Open weights | Yes | No |
| Input, per 1M tokens | $0.15 | $0.10 |
| Cache hit, per 1M | $0.03 | $0.01 |
| Cache write, per 1M | $0.15 (same as input) | $0.125 |
| Output, per 1M tokens | $0.50 | $0.50 |
| Runs in | OpenCode and OpenRouter | Codex, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. GPT-6 Luna: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | GLM-5.3-Flash | GPT-6 Luna |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.16 | $0.11 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.03 | $0.02 |
| Output-heavy generation, 30K input, 80K output | $0.04 | $0.04 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $17.60 | $11.55 |
| Where the session’s cost goes | ||
| Cache writes | $0.06 | $0.05 |
| Cache reads | $0.06 | $0.02 |
| Uncached input | $0.02 | $0.01 |
| Output | $0.03 | $0.03 |
- caching saves on the session with GLM-5.3-Flash (60%)
- $0.24
- caching saves on the session with GPT-6 Luna (61%)
- $0.17
Same output price, different cache discounts
GPT-6 Luna and GLM-5.3-Flash both charge $0.50 per million output tokens, so the output-heavy generation costs $0.04 on each. Input is 1.5x apart, $0.10 on Luna and $0.15 on GLM-5.3-Flash, and the large one-off review follows it at $0.02 against $0.03.
The session gap comes mostly from the cache. OpenAI bills a hit at 10% of input, $0.01 on Luna. Z.ai bills $0.03, 20% of GLM-5.3-Flash's input price. The example session reads 2M tokens from the cache, which costs $0.02 on Luna and $0.06 on GLM-5.3-Flash, and that line alone is $0.04 of the $0.05 gap.
Cache writes almost cancel out. OpenAI charges 1.25x input to write the cache from GPT-5.6 on, $0.125 per million on Luna, and Z.ai lists no write fee, so GLM-5.3-Flash's written tokens cost its $0.15 input rate. Even with OpenAI's premium, Luna's writes come in lower: $0.05 against $0.06 for the session's 400K written tokens.
What $0.05 a session adds up to
At 110 sessions a month, five each working day, the example comes to $11.55 on GPT-6 Luna and $17.60 on GLM-5.3-Flash. That is 52% more on GLM-5.3-Flash, but in dollars it is $6.05 a month.
Caching does similar work on both. It saves 61% of the uncached session cost on Luna and 60% on GLM-5.3-Flash: Z.ai's shallower hit discount is offset by writes that carry no premium, and OpenAI's deeper discount is offset by the 1.25x write price.
Rates on OpenRouter can differ from both lists. It sends a request for an open-weight model like GLM-5.3-Flash to one of several providers, each with its own price. The tables here use Z.ai's and OpenAI's own API prices.
How OpenAI and Z.ai position these models
In OpenAI's words, GPT-6 Luna is "our most efficient model for focused, high-volume tasks." It is the model Codex's documentation points to for narrow jobs that repeat. Its reasoning effort starts at medium, and Codex lets it go up to max. Codex, OpenAI's own agent, offers it, though not in Codex cloud, and Free and Go plans get it in the Codex app.
Z.ai describes GLM-5.3-Flash as a low-cost, natively multimodal model that outperforms GLM-5.2 at one-tenth the price. It highlights visual coding, where the model looks at interfaces and rendered results to test and improve its work, and hybrid sparse and linear attention that shrinks attention compute and KV cache size. The weights are MIT-licensed, so self-hosting is an option this page does not price.
Both have large windows: 1.05M tokens of context on Luna and 1M on GLM-5.3-Flash, with 128K of output on each. Luna has a long-context rule worth knowing if your agent loads whole repositories: past 272K input tokens, OpenAI bills the entire request at 2x for input and cache and 1.5x for output.
EveryToken prices GPT-6 Luna at OpenAI's rates in Codex and OpenCode, and prices GLM-5.3-Flash when you use it through OpenRouter.
Prompt caching
How each maker bills cached tokens
Z.ai
Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.
A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.
Source: Z.ai docs: Context caching
OpenAI
Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.
From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.
On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.
Source: OpenAI: Prompt caching
Your own numbers
See what GLM-5.3-Flash and GPT-6 Luna really cost you.
everyaitoken reads your OpenRouter, Codex, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is GPT-6 Luna cheaper than GLM-5.3-Flash?
At list prices, yes, on input and cache hits, while output costs $0.50 per million on both. The example cached session comes to $0.11 on GPT-6 Luna and $0.16 on GLM-5.3-Flash.
Why does GLM-5.3-Flash cost more on a cached session?
Z.ai prices a cache hit at 20% of input, $0.03 per million, while OpenAI prices it at 10%, $0.01 on Luna. Reading the session's 2M cached tokens costs $0.06 against $0.02.
Where can I use GLM-5.3-Flash?
Through OpenRouter and OpenCode, or on Z.ai's API. It is not in Cursor or GitHub Copilot, and its MIT-licensed weights also allow self-hosting.
What happens above 272K input tokens on GPT-6 Luna?
OpenAI switches the whole request to 2x for input and cache and 1.5x for output. Each request in the example session stays below 200K, so the tables use Luna's standard rates.
Sources
- Z.ai docs: Pricing
- Z.ai docs: GLM-5.3-Flash
- Z.ai: GLM-5.3-Flash
- OpenRouter: GLM-5.3-Flash
- OpenRouter: Rankings
- OpenCode docs: Zen
- OpenAI: API pricing
- OpenAI docs: GPT-6 Luna
- OpenAI: Introducing GPT-6 Sol and Luna
- OpenAI: API changelog
- Codex docs: Models
- OpenRouter: GPT-6 Luna
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Z.ai docs: Context caching
- OpenAI: Prompt caching