Model comparison
MiniMax M3 vs GLM-5.3-Flash: open-weight pricing compared
GLM-5.3-Flash costs about half as much as MiniMax M3 on every line, $0.16 against $0.33 per cached coding session. Output limits and licenses differ too.
· Prices as of September 28, 2026
MiniMax M3
MiniMax · Released June 1, 2026
MiniMax's open-weight model pitched against closed frontier models, with a 1M context and native multimodality.
MiniMax M3 facts and comparisonsGLM-5.3-Flash
Z.ai · Released August 26, 2026
Z.ai's low-cost, natively multimodal model, which Z.ai says outperforms GLM-5.2 at a tenth of the price.
GLM-5.3-Flash facts and comparisons
The short answer
GLM-5.3-Flash is the cheaper of these two open-weight models on every line, and the example agentic coding session costs $0.16 on it against $0.33 on MiniMax M3. MiniMax M3 makes sense if you need outputs beyond 128K tokens, up to 524.3K, or want the video input and desktop computer use MiniMax lists among its strengths. Both makers price a cache hit at 20% of input and charge nothing to write the cache, so the 2.1x session gap tracks their list prices.
Choose MiniMax M3 if
- You need long single responses: MiniMax lists 524.3K output tokens, 4.1x GLM-5.3-Flash's 128K, and recommends up to 131,072 per request.
- Your agent needs video input or desktop computer use, both of which MiniMax lists among M3's strengths.
- MiniMax's pitch fits your work: it says M3 reaches frontier-level performance on specialized tasks such as coding and agentic work.
Choose GLM-5.3-Flash if
- You want the lower price: $0.15 input and $0.50 output per million, against $0.30 and $1.20.
- Your work is output-heavy, and the generation workload costs $0.04 on GLM-5.3-Flash against $0.11.
- You want the MIT license rather than MiniMax's community license.
- You want a model that inspects rendered interfaces while it codes, which Z.ai calls visual coding.
Side by side
Specs and prices
| Fact | MiniMax M3 | GLM-5.3-Flash |
|---|---|---|
| Maker | MiniMax | Z.ai |
| API model id | MiniMax-M3 | glm-5.3-flash |
| Released | June 1, 2026 | August 26, 2026 |
| Status | Current | Current |
| Context window | 1M tokens | 1M tokens |
| Max output | 524.3K tokens | 128K tokens |
| Open weights | Yes | Yes |
| Input, per 1M tokens | $0.30 | $0.15 |
| Cache hit, per 1M | $0.06 | $0.03 |
| Cache write, per 1M | $0.30 (same as input) | $0.15 (same as input) |
| Output, per 1M tokens | $1.20 | $0.50 |
| Runs in | OpenCode and OpenRouter | OpenCode and OpenRouter |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. MiniMax M3: MiniMax labels these rates a permanent 50% discount. Requests over 512K input tokens cost $0.60 input, $0.12 cached, and $2.40 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | MiniMax M3 | GLM-5.3-Flash |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.33 | $0.16 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.06 | $0.03 |
| Output-heavy generation, 30K input, 80K output | $0.11 | $0.04 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $36.30 | $17.60 |
| Where the session’s cost goes | ||
| Cache writes | $0.12 | $0.06 |
| Cache reads | $0.12 | $0.06 |
| Uncached input | $0.03 | $0.02 |
| Output | $0.06 | $0.03 |
- caching saves on the session with MiniMax M3 (59%)
- $0.48
- caching saves on the session with GLM-5.3-Flash (60%)
- $0.24
Why the MiniMax M3 and GLM-5.3-Flash gap tracks list prices
MiniMax M3 charges $0.30 per million input tokens and $1.20 per million output tokens. GLM-5.3-Flash charges $0.15 and $0.50 on Z.ai's API. That is 2x on input and 2.4x on output, so the large one-off review costs $0.06 against $0.03, and the output-heavy generation $0.11 against $0.04, 2.8x.
Caching barely moves the ratio, because both makers price it the same way. Neither charges a fee to write the cache, and both bill a hit at 20% of input: $0.06 per million on MiniMax M3 and $0.03 on GLM-5.3-Flash. In the example session, writes and reads each cost $0.12 on MiniMax and $0.06 on GLM-5.3-Flash.
The session totals $0.33 against $0.16, 2.1x, and 110 sessions a month come to $36.30 against $17.60, an $18.70 difference. Caching saves 59% of the uncached cost on MiniMax and 60% on GLM-5.3-Flash, nearly the same share, since the discount rules match.
What MiniMax M3 offers for the higher price
The clearest spec difference is output length. MiniMax M3 lists 524.3K output tokens per response, 4.1x GLM-5.3-Flash's 128K, although MiniMax itself recommends staying at or under 131,072 output tokens per request. Both accept 1M tokens of context, and MiniMax attributes its window to MiniMax Sparse Attention.
MiniMax also lists image and video input and operating a desktop computer among M3's strengths, and describes M3 as reaching "frontier-level performance on specialized tasks such as coding and agentic work." Its current rates carry a label: MiniMax calls them a permanent 50% discount.
MiniMax has a long-context rule as well: requests over 512K input tokens cost $0.60 input, $0.12 cached, and $2.40 output per million. Coding sessions that stay under that line, like the example, pay the standard rates. MiniMax caches automatically once a request reaches 512 input tokens.
GLM-5.3-Flash: Z.ai's low-cost multimodal model
Z.ai positions GLM-5.3-Flash as a low-cost, natively multimodal model and says it outperforms GLM-5.2 at a tenth of the price. It points to visual coding, where the model looks at interfaces and rendered results to test and improve its work, and to a hybrid of sparse and linear attention that reduces attention compute and KV cache size.
Summed across nine languages, GLM-5.3-Flash was the most-used model for programming on OpenRouter over the week before September 28, 2026. On the caching side, Z.ai adds that storing GLM-5.3-Flash's cached input costs nothing for a limited time.
Licenses differ: GLM-5.3-Flash is MIT-licensed, and MiniMax M3 uses the MiniMax community license. Both can be self-hosted, and hosting costs are not estimated here. Both also run in OpenCode, and through OpenRouter, where a request goes to one of several providers whose prices can differ from the makers' own API prices in the tables.
EveryToken prices MiniMax M3 and GLM-5.3-Flash when they run through OpenRouter, from OpenRouter's catalog.
Prompt caching
How each maker bills cached tokens
MiniMax
MiniMax M3 caches repeated prompts automatically for requests of 512 or more input tokens, with no charge for cache writes.
A cache hit costs $0.06 per million tokens, one-fifth of the input price.
Source: MiniMax docs: Prompt caching
Z.ai
Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.
A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.
Source: Z.ai docs: Context caching
Your own numbers
See what MiniMax M3 and GLM-5.3-Flash really cost you.
everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is GLM-5.3-Flash cheaper than MiniMax M3?
Yes, on every line: $0.15 against $0.30 for input, $0.50 against $1.20 for output, and $0.03 against $0.06 for a cache hit. The example cached session costs $0.16 against $0.33.
Do MiniMax and Z.ai charge to write the cache?
No. Both cache automatically, list no charge for cache writes, and price a hit at 20% of input. MiniMax caches requests of 512 or more input tokens.
Which model has the larger output limit?
MiniMax M3, at 524.3K tokens per response against 128K for GLM-5.3-Flash. MiniMax recommends up to 131,072 output tokens per request.
Are MiniMax M3 and GLM-5.3-Flash open source?
Both publish open weights. GLM-5.3-Flash uses the MIT license and MiniMax M3 the MiniMax community license, so check the terms before building on either.