Skip to content

Model comparison

MiniMax M3 vs GPT-6 Luna: cents per session, 3x apart

GPT-6 Luna costs $0.11 per cached coding session against $0.33 on MiniMax M3. Where the 3x gap comes from, and what MiniMax M3 offers in return.

· Prices as of September 28, 2026

  • MiniMax M3

    MiniMax · Released June 1, 2026

    MiniMax's open-weight model pitched against closed frontier models, with a 1M context and native multimodality.

    MiniMax M3 facts and comparisons
  • GPT-6 Luna

    OpenAI · Released September 22, 2026

    The cheapest GPT-6 model, pitched for focused, high-volume, repeatable work, including narrower coding tasks.

    GPT-6 Luna facts and comparisons

The short answer

GPT-6 Luna costs about a third as much as MiniMax M3 on every example workload: $0.11 against $0.33 for the agentic coding session, and $11.55 against $36.30 for a month of 110 sessions. The biggest single difference is the cache hit, $0.01 per million on Luna against $0.06 on MiniMax M3. MiniMax M3 offers open weights, up to 524.3K tokens of output, and a long-context line at 512K rather than Luna's 272K, while Luna runs in Codex and GitHub Copilot.

Choose MiniMax M3 if

  • You want open weights, published under the MiniMax community license.
  • You want image and video input or desktop computer use, which MiniMax lists among M3's capabilities.
  • You need long outputs: MiniMax lists 524.3K as the maximum, against 128K on Luna.

Choose GPT-6 Luna if

  • Cost per token decides: $0.10 input, $0.01 cached, and $0.50 output per million.
  • You work in Codex, or in GitHub Copilot, which offers Luna and does not list MiniMax M3.
  • Your tasks are focused and repeatable, which is what OpenAI and the Codex docs recommend Luna for.

Side by side

Specs and prices

FactMiniMax M3GPT-6 Luna
MakerMiniMaxOpenAI
API model idMiniMax-M3gpt-6-luna
ReleasedJune 1, 2026September 22, 2026
StatusCurrentCurrent
Context window1M tokens1.05M tokens
Max output524.3K tokens128K tokens
Open weightsYesNo
Input, per 1M tokens$0.30$0.10
Cache hit, per 1M$0.06$0.01
Cache write, per 1M$0.30 (same as input)$0.125
Output, per 1M tokens$1.20$0.50
Runs inOpenCode and OpenRouterCodex, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. MiniMax M3: MiniMax labels these rates a permanent 50% discount. Requests over 512K input tokens cost $0.60 input, $0.12 cached, and $2.40 output per million tokens. GPT-6 Luna: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadMiniMax M3GPT-6 Luna
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$0.33$0.11
Large one-off review, 150K input with no cache hits, 10K output$0.06$0.02
Output-heavy generation, 30K input, 80K output$0.11$0.04
A month of sessions, 110 sessions: 5 a day, 22 working days$36.30$11.55
Where the session’s cost goes
Cache writes$0.12$0.05
Cache reads$0.12$0.02
Uncached input$0.03$0.01
Output$0.06$0.03
caching saves on the session with MiniMax M3 (59%)
$0.48
caching saves on the session with GPT-6 Luna (61%)
$0.17

GPT-6 Luna is cheaper on every line of the rate card

Per million tokens, GPT-6 Luna asks $0.10 for input, $0.01 for a cache hit, $0.125 for a cache write, and $0.50 for output, where MiniMax M3 asks $0.30, $0.06, $0.30, and $1.20. Input is 3x dearer on MiniMax M3, output 2.4x, and cache hits 6x.

OpenAI adds a 1.25x premium to cache writes and MiniMax adds none, yet Luna's input is low enough that its premium write still costs less than a MiniMax M3 write billed as plain input. In the example session, writes cost $0.05 on Luna and $0.12 on MiniMax M3, and reads cost $0.02 and $0.12.

The session totals $0.11 against $0.33, a 3x gap. The large one-off review shows the same ratio, $0.02 against $0.06, and the output-heavy generation a slightly smaller one, 2.8x, at $0.04 against $0.11. Across 110 sessions a month the two differ by $24.75.

Where MiniMax M3's limits run further

The context windows are close: 1M on MiniMax M3 and 1.05M on Luna. The long-context rules are not. Luna switches to 2x input and cache pricing and 1.5x output pricing once a request carries more than 272K input tokens, and the switch applies to the entire request. MiniMax M3 keeps its standard rates up to 512K input tokens, and above that charges $0.60 input, $0.12 cached, and $2.40 output per million.

So between 272K and 512K, Luna's rates rise and MiniMax M3's do not. Luna still costs less per token in that band, but the gap narrows. Above 512K both are on their higher tiers.

Output limits differ more clearly. MiniMax lists a 524.3K maximum output for M3 and recommends up to 131,072 tokens per request. Luna writes up to 128K per request.

What MiniMax and OpenAI say about these models

MiniMax pitches M3 against closed frontier models: "M3 reaches frontier-level performance on specialized tasks such as coding and agentic work." OpenAI calls Luna "our most efficient model for focused, high-volume tasks," the cheapest model of its GPT-6 family, and the Codex docs recommend it for focused, repeatable tasks. The two makers aim at different jobs as much as at different prices.

Luna runs in Codex, OpenAI's own coding agent, though not in Codex cloud, and Free and Go plans get it in the Codex app. It is also in OpenCode, OpenRouter, and GitHub Copilot. MiniMax M3, whose weights MiniMax releases under the MiniMax community license, is available here through OpenCode and OpenRouter. On OpenRouter an open-weight model is served by one of several providers, whose prices can differ from MiniMax's own API, and the tables use MiniMax's price.

EveryToken prices Luna at OpenAI's rates from Codex or OpenCode history, and prices MiniMax M3 through OpenRouter, from OpenRouter's catalog. For two models this inexpensive, the more useful view is often how much of each session the cache covered, which it shows per model.

Prompt caching

How each maker bills cached tokens

MiniMax

MiniMax M3 caches repeated prompts automatically for requests of 512 or more input tokens, with no charge for cache writes.

A cache hit costs $0.06 per million tokens, one-fifth of the input price.

Source: MiniMax docs: Prompt caching

OpenAI

Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.

From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.

On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.

Source: OpenAI: Prompt caching

Your own numbers

See what MiniMax M3 and GPT-6 Luna really cost you.

everyaitoken reads your OpenRouter, Codex, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is GPT-6 Luna cheaper than MiniMax M3?

Yes, on every rate and every example workload. The agentic coding session costs $0.11 on Luna against $0.33 on MiniMax M3, and output costs $0.50 per million against $1.20.

Where do the long-context price lines fall on each model?

On GPT-6 Luna the line sits at 272K input tokens, past which input and cache cost 2x and output 1.5x. MiniMax M3 keeps standard rates up to 512K input tokens, then charges $0.60 input, $0.12 cached, and $2.40 output per million.

Can I self-host MiniMax M3?

MiniMax publishes open weights under the MiniMax community license, so self-hosting is possible within its terms. GPT-6 Luna has no open weights. This post does not estimate hosting costs.

Does MiniMax charge for cache writes?

No. MiniMax M3 caches repeated prompts automatically for requests of 512 or more input tokens, with no charge for cache writes. OpenAI bills Luna's writes at 1.25x input, $0.125 per million.

  • GPT-6 Astra vs GPT-6 Luna

    GPT-6 Astra and GPT-6 Luna sit at opposite ends of OpenAI's lineup, 100x apart per token. What each is for, and where GPT-6 Sol fits between them.

  • GPT-6 Sol vs GPT-6 Luna

    GPT-6 Luna costs a twentieth of GPT-6 Sol per token. What OpenAI and the Codex docs say each tier is for, and what the gap means for coding sessions.

  • GPT-6 Luna vs GPT-5.6 Luna

    GPT-6 Luna halves the input price of GPT-5.6 Luna, which already had an 80% cut, and trims output further. What the move saves, and what Codex suggests.

  • Claude Haiku 4.5 vs GPT-6 Luna

    GPT-6 Luna lists a tenth of Claude Haiku 4.5's rates and a 1.05M context window against 200K. What each costs for sub-agents and cached coding sessions.

  • DeepSeek-V4.1-Flash vs GPT-6 Luna

    GPT-6 Luna undercuts DeepSeek-V4.1-Flash at list prices, $0.11 against $0.22 for a cached coding session. How cache pricing and off-peak hours move the gap.

  • DeepSeek-V4.1-Flash vs MiniMax M3

    DeepSeek-V4.1-Flash and MiniMax M3 list the same $0.30 input and $1.20 output rates, but MiniMax's cache hits cost 10x as much: $0.33 vs $0.22 a session.