Skip to content

MiniMax

MiniMax M3: price, context window, and caching

MiniMax's open-weight model pitched against closed frontier models, with a 1M context and native multimodality.

Released June 1, 2026 · Prices as of September 28, 2026

In MiniMax’s words

“M3 reaches frontier-level performance on specialized tasks such as coding and agentic work.”

MiniMax: MiniMax M3

What MiniMax says it’s good at

  • A 1M context through MiniMax Sparse Attention Source
  • Image and video input, and operating a desktop computer Source

Facts

Specs and prices

FactMiniMax M3
MakerMiniMax
API model idMiniMax-M3
ReleasedJune 1, 2026
StatusCurrent
Context window1M tokens
Max output524.3K tokens
Open weightsYes
Input, per 1M tokens$0.30
Cache hit, per 1M$0.06
Cache write, per 1M$0.30 (same as input)
Output, per 1M tokens$1.20
Runs inOpenCode and OpenRouter

Standard API rates in US dollars, as published by the maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. MiniMax M3: MiniMax labels these rates a permanent 50% discount. Requests over 512K input tokens cost $0.60 input, $0.12 cached, and $2.40 output per million tokens.

Good to know

  • Open weights under the MiniMax community license.
  • MiniMax recommends up to 131,072 output tokens per request.

Cost

What typical work costs

Example token counts at MiniMax M3’s published rates. On the agentic session, caching saves $0.48 against billing every token as ordinary input.

Example workload costs for MiniMax M3
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$0.33
Large one-off review, 150K input with no cache hits, 10K output$0.06
Output-heavy generation, 30K input, 80K output$0.11
A month of sessions, 110 sessions: 5 a day, 22 working days$36.30

Prompt caching

How MiniMax bills cached tokens

MiniMax

MiniMax M3 caches repeated prompts automatically for requests of 512 or more input tokens, with no charge for cache writes.

A cache hit costs $0.06 per million tokens, one-fifth of the input price.

Source: MiniMax docs: Prompt caching

MiniMax M3 compared

  • DeepSeek-V4.1-Flash vs MiniMax M3

    DeepSeek-V4.1-Flash and MiniMax M3 list the same $0.30 input and $1.20 output rates, but MiniMax's cache hits cost 10x as much: $0.33 vs $0.22 a session.

  • MiniMax M3 vs Claude Sonnet 5

    MiniMax M3 costs $0.33 per cached coding session against $2.40 on Claude Sonnet 5. Both have a 1M window, and they differ on tools, caching, and long prompts.

  • MiniMax M3 vs Gemini 3.8 Flash

    MiniMax M3 costs $0.33 per cached coding session against $0.71 on Gemini 3.8 Flash, whose introductory rates end on December 31, 2026.

  • MiniMax M3 vs GLM-5.3-Flash

    GLM-5.3-Flash costs about half as much as MiniMax M3 on every line, $0.16 against $0.33 per cached coding session. Output limits and licenses differ too.

  • MiniMax M3 vs GPT-6 Luna

    GPT-6 Luna costs $0.11 per cached coding session against $0.33 on MiniMax M3. Where the 3x gap comes from, and what MiniMax M3 offers in return.

Your own numbers

See what MiniMax M3 really costs you.

everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math