Skip to content

Model comparison

MiniMax M3 vs Gemini 3.8 Flash: permanent discount vs promo

MiniMax M3 costs $0.33 per cached coding session against $0.71 on Gemini 3.8 Flash, whose introductory rates end on December 31, 2026.

· Prices as of September 28, 2026

  • MiniMax M3

    MiniMax · Released June 1, 2026

    MiniMax's open-weight model pitched against closed frontier models, with a 1M context and native multimodality.

    MiniMax M3 facts and comparisons
  • Gemini 3.8 Flash

    Google · Released September 2, 2026

    Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.

    Gemini 3.8 Flash facts and comparisons

The short answer

MiniMax M3 costs less than Gemini 3.8 Flash on every example workload, $0.33 against $0.71 for the agentic coding session and $0.11 against $0.32 for output-heavy work. The gap is smallest on cache reads, $0.06 against $0.075 per million, and it would widen at Flash's 2027 rates, since Google's introductory prices end on December 31, 2026, while MiniMax labels its own rates a permanent 50% discount. Gemini 3.8 Flash counters with Gemini CLI, Cursor, GitHub Copilot, and a free API tier, and MiniMax M3 with open weights and up to 524.3K tokens of output.

Choose MiniMax M3 if

  • Price decides: $0.30 input and $1.20 output per million, against $0.75 and $3.75.
  • You plan beyond 2026: MiniMax labels its rates a permanent 50% discount, while Flash's rates double from January 1, 2027.
  • You need long outputs, up to 524.3K tokens against Flash's 65.5K.
  • You want open weights under the MiniMax community license.

Choose Gemini 3.8 Flash if

  • You code in Gemini CLI, whose default auto model pairs Flash with a Pro model, or in Cursor or GitHub Copilot, which both list Flash.
  • You want a no-cost trial, and the Gemini API free tier extends to Flash's caching as well as its input and output.
  • Your long prompts are mostly cached: above 512K input tokens, a MiniMax M3 hit costs $0.12 against Flash's $0.075.
  • You want Google's pitch of long-horizon software engineering and better robustness against prompt injection.

Side by side

Specs and prices

FactMiniMax M3Gemini 3.8 Flash
MakerMiniMaxGoogle
API model idMiniMax-M3gemini-3.8-flash
ReleasedJune 1, 2026September 2, 2026
StatusCurrentCurrent
Context window1M tokens1.05M tokens
Max output524.3K tokens65.5K tokens
Open weightsYesNo
Input, per 1M tokens$0.30$0.75
Cache hit, per 1M$0.06$0.075
Cache write, per 1M$0.30 (same as input)$0.75 (same as input)
Output, per 1M tokens$1.20$3.75
Runs inOpenCode and OpenRouterCursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. MiniMax M3: MiniMax labels these rates a permanent 50% discount. Requests over 512K input tokens cost $0.60 input, $0.12 cached, and $2.40 output per million tokens. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadMiniMax M3Gemini 3.8 Flash
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$0.33$0.71
Large one-off review, 150K input with no cache hits, 10K output$0.06$0.15
Output-heavy generation, 30K input, 80K output$0.11$0.32
A month of sessions, 110 sessions: 5 a day, 22 working days$36.30$78.38
Where the session’s cost goes
Cache writes$0.12$0.30
Cache reads$0.12$0.15
Uncached input$0.03$0.08
Output$0.06$0.19
caching saves on the session with MiniMax M3 (59%)
$0.48
caching saves on the session with Gemini 3.8 Flash (66%)
$1.35

A permanent discount against an introductory price

MiniMax M3 lists $0.30 per million input tokens, $0.06 per cache hit, and $1.20 per million output tokens, and MiniMax labels these rates a permanent 50% discount. Gemini 3.8 Flash lists $0.75, $0.075, and $3.75, which Google calls introductory rates through December 31, 2026.

From January 1, 2027, Flash's rates become $1.50 input, $0.15 cached, and $7.50 output. Every comparison in the table therefore has an end date on the Flash side. If MiniMax's prices stay where they are, the 2.2x session gap shown today would roughly double at Flash's later rates.

Why cache reads narrow the gap

Most rates differ by 2.5x to 3.1x: input and cache writes cost 2.5x as much on Flash, and output 3.1x. Cache hits are the exception. MiniMax M3's hit costs one-fifth of its input price, and Flash's costs 10%, so the two land close together at $0.06 and $0.075.

That shapes the example session. Its 2M cached tokens cost $0.12 on MiniMax M3 and $0.15 on Flash, only $0.03 apart, while writes, $0.12 against $0.30, and output, $0.06 against $0.19, carry most of the difference. The session ends at $0.33 against $0.71, a 2.2x gap, narrower than the 2.5x on the uncached review, $0.06 against $0.15, and the 2.9x on the output-heavy generation, $0.11 against $0.32.

Neither maker charges a separate cache-write fee. MiniMax caches repeated prompts automatically once a request has 512 or more input tokens. Google's implicit caching is also automatic, though a hit is not assured, and its explicit caching adds storage at $0.50 to $1 per million tokens per hour on Flash models.

Context, output limits, and long prompts

Both accept about 1M tokens: 1M on MiniMax M3 and 1.05M on Flash. MiniMax M3 requests over 512K input tokens cost $0.60 input, $0.12 cached, and $2.40 output per million, and Flash's price notes list no long-context tier. Above 512K, a MiniMax M3 cache hit costs more than Flash's, while its input and output stay cheaper.

Output limits are far apart. Flash writes up to 65.5K tokens per request. MiniMax M3's maximum is 524.3K, and MiniMax recommends up to 131,072 per request, still about twice Flash's cap.

Tools, weights, and how each maker pitches its model

For Gemini API key and Vertex AI users, Gemini CLI, Google's own coding agent, routes its default auto model's Flash work to Gemini 3.8 Flash, and the model is also listed in Cursor, OpenCode, OpenRouter, and GitHub Copilot. Its default thinking level is medium, and Gemini CLI sends high. Of the tools here, MiniMax M3 is reachable through OpenCode and OpenRouter, and MiniMax releases its weights under the MiniMax community license. OpenRouter routes an open-weight model to one of several providers, whose prices can differ from MiniMax's own API, so the table uses MiniMax's price.

Google calls Flash "our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows." MiniMax says "M3 reaches frontier-level performance on specialized tasks such as coding and agentic work." EveryToken prices Flash at Google's rates from Gemini CLI, Cursor, or OpenCode history, and prices MiniMax M3 from OpenRouter's catalog when you use it there.

Prompt caching

How each maker bills cached tokens

MiniMax

MiniMax M3 caches repeated prompts automatically for requests of 512 or more input tokens, with no charge for cache writes.

A cache hit costs $0.06 per million tokens, one-fifth of the input price.

Source: MiniMax docs: Prompt caching

Google

Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.

Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.

On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.

Source: Google: Context caching

Your own numbers

See what MiniMax M3 and Gemini 3.8 Flash really cost you.

everyaitoken reads your OpenRouter, Gemini CLI, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is MiniMax M3 cheaper than Gemini 3.8 Flash?

Yes, on every example workload: $0.33 against $0.71 for the agentic coding session, and $36.30 against $78.38 for a month of 110 sessions. The narrowest gap is on cache hits, $0.06 against $0.075 per million.

When do Gemini 3.8 Flash's prices go up?

Its current rates are introductory through December 31, 2026. On January 1, 2027 that changes to $1.50 input, $0.15 cached, and $7.50 output per million, twice today's Flash rates.

Which model can write longer outputs?

MiniMax M3, with a 524.3K maximum and a recommended ceiling of 131,072 tokens per request. Gemini 3.8 Flash writes up to 65.5K tokens per request.

Are MiniMax M3's weights open?

Yes. MiniMax publishes them under the MiniMax community license. Gemini 3.8 Flash has no open weights, and this post does not estimate self-hosting costs.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash

    Gemini 3.8 Flash costs about half of Gemini 3.5 Flash on introductory rates. What changes on January 1, 2027, and which one Gemini CLI picks for you.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash-Lite

    Google recommends Gemini 3.8 Flash and Gemini 3.5 Flash-Lite together for new projects. Where each fits, the 2.1x session gap, and what changes in 2027.

  • Gemini 3.8 Flash vs Gemini 3.1 Pro Preview

    Gemini CLI's auto model uses both Gemini 3.8 Flash and Gemini 3.1 Pro Preview. What each costs, why Flash rates double in 2027, and the Pro's preview status.

  • Gemini 3.8 Flash vs Gemini 3.7 Flash

    Gemini 3.8 Flash and Gemini 3.7 Flash cost the same, and both rates double on January 1, 2027. What differs is Gemini CLI, Cursor, and Google's own pitch.

  • Claude Fable 5.1 vs Gemini 3.8 Flash

    A cached coding session costs 14.8x more on Claude Fable 5.1 than on Gemini 3.8 Flash. How Flash's introductory rates, output cap, and free tier factor in.

  • Claude Haiku 4.5 vs Gemini 3.8 Flash

    Gemini 3.8 Flash undercuts Claude Haiku 4.5 on introductory rates, $0.71 against $1.20 per cached session. From January 1, 2027, its rates top Haiku's.