Skip to content

Model comparison

MiniMax M3 vs GLM-5.3-Flash: open-weight pricing compared

GLM-5.3-Flash costs about half as much as MiniMax M3 on every line, $0.16 against $0.33 per cached coding session. Output limits and licenses differ too.

· Prices as of September 28, 2026

  • MiniMax M3

    MiniMax · Released June 1, 2026

    MiniMax's open-weight model pitched against closed frontier models, with a 1M context and native multimodality.

    MiniMax M3 facts and comparisons
  • GLM-5.3-Flash

    Z.ai · Released August 26, 2026

    Z.ai's low-cost, natively multimodal model, which Z.ai says outperforms GLM-5.2 at a tenth of the price.

    GLM-5.3-Flash facts and comparisons

The short answer

GLM-5.3-Flash is the cheaper of these two open-weight models on every line, and the example agentic coding session costs $0.16 on it against $0.33 on MiniMax M3. MiniMax M3 makes sense if you need outputs beyond 128K tokens, up to 524.3K, or want the video input and desktop computer use MiniMax lists among its strengths. Both makers price a cache hit at 20% of input and charge nothing to write the cache, so the 2.1x session gap tracks their list prices.

Choose MiniMax M3 if

  • You need long single responses: MiniMax lists 524.3K output tokens, 4.1x GLM-5.3-Flash's 128K, and recommends up to 131,072 per request.
  • Your agent needs video input or desktop computer use, both of which MiniMax lists among M3's strengths.
  • MiniMax's pitch fits your work: it says M3 reaches frontier-level performance on specialized tasks such as coding and agentic work.

Choose GLM-5.3-Flash if

  • You want the lower price: $0.15 input and $0.50 output per million, against $0.30 and $1.20.
  • Your work is output-heavy, and the generation workload costs $0.04 on GLM-5.3-Flash against $0.11.
  • You want the MIT license rather than MiniMax's community license.
  • You want a model that inspects rendered interfaces while it codes, which Z.ai calls visual coding.

Side by side

Specs and prices

FactMiniMax M3GLM-5.3-Flash
MakerMiniMaxZ.ai
API model idMiniMax-M3glm-5.3-flash
ReleasedJune 1, 2026August 26, 2026
StatusCurrentCurrent
Context window1M tokens1M tokens
Max output524.3K tokens128K tokens
Open weightsYesYes
Input, per 1M tokens$0.30$0.15
Cache hit, per 1M$0.06$0.03
Cache write, per 1M$0.30 (same as input)$0.15 (same as input)
Output, per 1M tokens$1.20$0.50
Runs inOpenCode and OpenRouterOpenCode and OpenRouter

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. MiniMax M3: MiniMax labels these rates a permanent 50% discount. Requests over 512K input tokens cost $0.60 input, $0.12 cached, and $2.40 output per million tokens.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadMiniMax M3GLM-5.3-Flash
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$0.33$0.16
Large one-off review, 150K input with no cache hits, 10K output$0.06$0.03
Output-heavy generation, 30K input, 80K output$0.11$0.04
A month of sessions, 110 sessions: 5 a day, 22 working days$36.30$17.60
Where the session’s cost goes
Cache writes$0.12$0.06
Cache reads$0.12$0.06
Uncached input$0.03$0.02
Output$0.06$0.03
caching saves on the session with MiniMax M3 (59%)
$0.48
caching saves on the session with GLM-5.3-Flash (60%)
$0.24

Why the MiniMax M3 and GLM-5.3-Flash gap tracks list prices

MiniMax M3 charges $0.30 per million input tokens and $1.20 per million output tokens. GLM-5.3-Flash charges $0.15 and $0.50 on Z.ai's API. That is 2x on input and 2.4x on output, so the large one-off review costs $0.06 against $0.03, and the output-heavy generation $0.11 against $0.04, 2.8x.

Caching barely moves the ratio, because both makers price it the same way. Neither charges a fee to write the cache, and both bill a hit at 20% of input: $0.06 per million on MiniMax M3 and $0.03 on GLM-5.3-Flash. In the example session, writes and reads each cost $0.12 on MiniMax and $0.06 on GLM-5.3-Flash.

The session totals $0.33 against $0.16, 2.1x, and 110 sessions a month come to $36.30 against $17.60, an $18.70 difference. Caching saves 59% of the uncached cost on MiniMax and 60% on GLM-5.3-Flash, nearly the same share, since the discount rules match.

What MiniMax M3 offers for the higher price

The clearest spec difference is output length. MiniMax M3 lists 524.3K output tokens per response, 4.1x GLM-5.3-Flash's 128K, although MiniMax itself recommends staying at or under 131,072 output tokens per request. Both accept 1M tokens of context, and MiniMax attributes its window to MiniMax Sparse Attention.

MiniMax also lists image and video input and operating a desktop computer among M3's strengths, and describes M3 as reaching "frontier-level performance on specialized tasks such as coding and agentic work." Its current rates carry a label: MiniMax calls them a permanent 50% discount.

MiniMax has a long-context rule as well: requests over 512K input tokens cost $0.60 input, $0.12 cached, and $2.40 output per million. Coding sessions that stay under that line, like the example, pay the standard rates. MiniMax caches automatically once a request reaches 512 input tokens.

GLM-5.3-Flash: Z.ai's low-cost multimodal model

Z.ai positions GLM-5.3-Flash as a low-cost, natively multimodal model and says it outperforms GLM-5.2 at a tenth of the price. It points to visual coding, where the model looks at interfaces and rendered results to test and improve its work, and to a hybrid of sparse and linear attention that reduces attention compute and KV cache size.

Summed across nine languages, GLM-5.3-Flash was the most-used model for programming on OpenRouter over the week before September 28, 2026. On the caching side, Z.ai adds that storing GLM-5.3-Flash's cached input costs nothing for a limited time.

Licenses differ: GLM-5.3-Flash is MIT-licensed, and MiniMax M3 uses the MiniMax community license. Both can be self-hosted, and hosting costs are not estimated here. Both also run in OpenCode, and through OpenRouter, where a request goes to one of several providers whose prices can differ from the makers' own API prices in the tables.

EveryToken prices MiniMax M3 and GLM-5.3-Flash when they run through OpenRouter, from OpenRouter's catalog.

Prompt caching

How each maker bills cached tokens

MiniMax

MiniMax M3 caches repeated prompts automatically for requests of 512 or more input tokens, with no charge for cache writes.

A cache hit costs $0.06 per million tokens, one-fifth of the input price.

Source: MiniMax docs: Prompt caching

Z.ai

Z.ai caches repeated context automatically, with no configuration, and lists no fee for writing the cache.

A cache hit costs $0.26 per million tokens on GLM-5.3 and $0.03 on GLM-5.3-Flash, and cached input storage is free for a limited time.

Source: Z.ai docs: Context caching

Your own numbers

See what MiniMax M3 and GLM-5.3-Flash really cost you.

everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is GLM-5.3-Flash cheaper than MiniMax M3?

Yes, on every line: $0.15 against $0.30 for input, $0.50 against $1.20 for output, and $0.03 against $0.06 for a cache hit. The example cached session costs $0.16 against $0.33.

Do MiniMax and Z.ai charge to write the cache?

No. Both cache automatically, list no charge for cache writes, and price a hit at 20% of input. MiniMax caches requests of 512 or more input tokens.

Which model has the larger output limit?

MiniMax M3, at 524.3K tokens per response against 128K for GLM-5.3-Flash. MiniMax recommends up to 131,072 output tokens per request.

Are MiniMax M3 and GLM-5.3-Flash open source?

Both publish open weights. GLM-5.3-Flash uses the MIT license and MiniMax M3 the MiniMax community license, so check the terms before building on either.

  • GLM-5.3 vs GLM-5.3-Flash

    GLM-5.3-Flash costs about a ninth of GLM-5.3 at Z.ai's rates: $0.16 against $1.44 for an agentic coding session. Same limits, same caching, different jobs.

  • GLM-5.3-Flash vs Claude Sonnet 5

    GLM-5.3-Flash costs $0.16 for a cached coding session that costs $2.40 on Claude Sonnet 5. What drives a 15x gap, and what price alone can't tell you.

  • DeepSeek-V4.1-Flash vs GLM-5.3-Flash

    GLM-5.3-Flash lists half DeepSeek-V4.1-Flash's input price, but GLM's cache hits cost 5x as much. A cached coding session: $0.16 against $0.22 at peak.

  • DeepSeek-V4.1-Flash vs MiniMax M3

    DeepSeek-V4.1-Flash and MiniMax M3 list the same $0.30 input and $1.20 output rates, but MiniMax's cache hits cost 10x as much: $0.33 vs $0.22 a session.

  • GLM-5.3-Flash vs Claude Haiku 4.5

    GLM-5.3-Flash charges $0.50 per million output tokens against $5 on Claude Haiku 4.5, and $0.16 against $1.20 for a cached coding session. Where each fits.

  • GLM-5.3-Flash vs Gemini 3.8 Flash

    GLM-5.3-Flash costs $0.16 for a cached coding session against $0.71 on Gemini 3.8 Flash, whose introductory rates end December 31, 2026. What else differs.