Model comparison
MiniMax M3 vs Gemini 3.8 Flash: permanent discount vs promo
MiniMax M3 costs $0.33 per cached coding session against $0.71 on Gemini 3.8 Flash, whose introductory rates end on December 31, 2026.
· Prices as of September 28, 2026
MiniMax M3
MiniMax · Released June 1, 2026
MiniMax's open-weight model pitched against closed frontier models, with a 1M context and native multimodality.
MiniMax M3 facts and comparisonsGemini 3.8 Flash
Google · Released September 2, 2026
Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.
Gemini 3.8 Flash facts and comparisons
The short answer
MiniMax M3 costs less than Gemini 3.8 Flash on every example workload, $0.33 against $0.71 for the agentic coding session and $0.11 against $0.32 for output-heavy work. The gap is smallest on cache reads, $0.06 against $0.075 per million, and it would widen at Flash's 2027 rates, since Google's introductory prices end on December 31, 2026, while MiniMax labels its own rates a permanent 50% discount. Gemini 3.8 Flash counters with Gemini CLI, Cursor, GitHub Copilot, and a free API tier, and MiniMax M3 with open weights and up to 524.3K tokens of output.
Choose MiniMax M3 if
- Price decides: $0.30 input and $1.20 output per million, against $0.75 and $3.75.
- You plan beyond 2026: MiniMax labels its rates a permanent 50% discount, while Flash's rates double from January 1, 2027.
- You need long outputs, up to 524.3K tokens against Flash's 65.5K.
- You want open weights under the MiniMax community license.
Choose Gemini 3.8 Flash if
- You code in Gemini CLI, whose default auto model pairs Flash with a Pro model, or in Cursor or GitHub Copilot, which both list Flash.
- You want a no-cost trial, and the Gemini API free tier extends to Flash's caching as well as its input and output.
- Your long prompts are mostly cached: above 512K input tokens, a MiniMax M3 hit costs $0.12 against Flash's $0.075.
- You want Google's pitch of long-horizon software engineering and better robustness against prompt injection.
Side by side
Specs and prices
| Fact | MiniMax M3 | Gemini 3.8 Flash |
|---|---|---|
| Maker | MiniMax | |
| API model id | MiniMax-M3 | gemini-3.8-flash |
| Released | June 1, 2026 | September 2, 2026 |
| Status | Current | Current |
| Context window | 1M tokens | 1.05M tokens |
| Max output | 524.3K tokens | 65.5K tokens |
| Open weights | Yes | No |
| Input, per 1M tokens | $0.30 | $0.75 |
| Cache hit, per 1M | $0.06 | $0.075 |
| Cache write, per 1M | $0.30 (same as input) | $0.75 (same as input) |
| Output, per 1M tokens | $1.20 | $3.75 |
| Runs in | OpenCode and OpenRouter | Cursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. MiniMax M3: MiniMax labels these rates a permanent 50% discount. Requests over 512K input tokens cost $0.60 input, $0.12 cached, and $2.40 output per million tokens. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | MiniMax M3 | Gemini 3.8 Flash |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.33 | $0.71 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.06 | $0.15 |
| Output-heavy generation, 30K input, 80K output | $0.11 | $0.32 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $36.30 | $78.38 |
| Where the session’s cost goes | ||
| Cache writes | $0.12 | $0.30 |
| Cache reads | $0.12 | $0.15 |
| Uncached input | $0.03 | $0.08 |
| Output | $0.06 | $0.19 |
- caching saves on the session with MiniMax M3 (59%)
- $0.48
- caching saves on the session with Gemini 3.8 Flash (66%)
- $1.35
A permanent discount against an introductory price
MiniMax M3 lists $0.30 per million input tokens, $0.06 per cache hit, and $1.20 per million output tokens, and MiniMax labels these rates a permanent 50% discount. Gemini 3.8 Flash lists $0.75, $0.075, and $3.75, which Google calls introductory rates through December 31, 2026.
From January 1, 2027, Flash's rates become $1.50 input, $0.15 cached, and $7.50 output. Every comparison in the table therefore has an end date on the Flash side. If MiniMax's prices stay where they are, the 2.2x session gap shown today would roughly double at Flash's later rates.
Why cache reads narrow the gap
Most rates differ by 2.5x to 3.1x: input and cache writes cost 2.5x as much on Flash, and output 3.1x. Cache hits are the exception. MiniMax M3's hit costs one-fifth of its input price, and Flash's costs 10%, so the two land close together at $0.06 and $0.075.
That shapes the example session. Its 2M cached tokens cost $0.12 on MiniMax M3 and $0.15 on Flash, only $0.03 apart, while writes, $0.12 against $0.30, and output, $0.06 against $0.19, carry most of the difference. The session ends at $0.33 against $0.71, a 2.2x gap, narrower than the 2.5x on the uncached review, $0.06 against $0.15, and the 2.9x on the output-heavy generation, $0.11 against $0.32.
Neither maker charges a separate cache-write fee. MiniMax caches repeated prompts automatically once a request has 512 or more input tokens. Google's implicit caching is also automatic, though a hit is not assured, and its explicit caching adds storage at $0.50 to $1 per million tokens per hour on Flash models.
Context, output limits, and long prompts
Both accept about 1M tokens: 1M on MiniMax M3 and 1.05M on Flash. MiniMax M3 requests over 512K input tokens cost $0.60 input, $0.12 cached, and $2.40 output per million, and Flash's price notes list no long-context tier. Above 512K, a MiniMax M3 cache hit costs more than Flash's, while its input and output stay cheaper.
Output limits are far apart. Flash writes up to 65.5K tokens per request. MiniMax M3's maximum is 524.3K, and MiniMax recommends up to 131,072 per request, still about twice Flash's cap.
Tools, weights, and how each maker pitches its model
For Gemini API key and Vertex AI users, Gemini CLI, Google's own coding agent, routes its default auto model's Flash work to Gemini 3.8 Flash, and the model is also listed in Cursor, OpenCode, OpenRouter, and GitHub Copilot. Its default thinking level is medium, and Gemini CLI sends high. Of the tools here, MiniMax M3 is reachable through OpenCode and OpenRouter, and MiniMax releases its weights under the MiniMax community license. OpenRouter routes an open-weight model to one of several providers, whose prices can differ from MiniMax's own API, so the table uses MiniMax's price.
Google calls Flash "our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows." MiniMax says "M3 reaches frontier-level performance on specialized tasks such as coding and agentic work." EveryToken prices Flash at Google's rates from Gemini CLI, Cursor, or OpenCode history, and prices MiniMax M3 from OpenRouter's catalog when you use it there.
Prompt caching
How each maker bills cached tokens
MiniMax
MiniMax M3 caches repeated prompts automatically for requests of 512 or more input tokens, with no charge for cache writes.
A cache hit costs $0.06 per million tokens, one-fifth of the input price.
Source: MiniMax docs: Prompt caching
Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.
Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.
On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.
Source: Google: Context caching
Your own numbers
See what MiniMax M3 and Gemini 3.8 Flash really cost you.
everyaitoken reads your OpenRouter, Gemini CLI, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is MiniMax M3 cheaper than Gemini 3.8 Flash?
Yes, on every example workload: $0.33 against $0.71 for the agentic coding session, and $36.30 against $78.38 for a month of 110 sessions. The narrowest gap is on cache hits, $0.06 against $0.075 per million.
When do Gemini 3.8 Flash's prices go up?
Its current rates are introductory through December 31, 2026. On January 1, 2027 that changes to $1.50 input, $0.15 cached, and $7.50 output per million, twice today's Flash rates.
Which model can write longer outputs?
MiniMax M3, with a 524.3K maximum and a recommended ceiling of 131,072 tokens per request. Gemini 3.8 Flash writes up to 65.5K tokens per request.
Are MiniMax M3's weights open?
Yes. MiniMax publishes them under the MiniMax community license. Gemini 3.8 Flash has no open weights, and this post does not estimate self-hosting costs.
Sources
- MiniMax docs: Pay-as-you-go pricing
- MiniMax: MiniMax M3
- OpenRouter: MiniMax M3
- OpenCode docs: Zen
- Google: Gemini API pricing
- Google docs: Gemini 3.8 Flash
- Google: Gemini models
- Google: Latest Gemini model
- Google DeepMind: Gemini 3.8 Flash model card
- Google: Gemini 3.8 Flash
- Gemini CLI source: model configuration
- Cursor docs: Gemini 3.8 Flash
- OpenRouter: Gemini 3.8 Flash
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- MiniMax docs: Prompt caching
- Google: Context caching