Model comparison
MiniMax M3 vs Claude Sonnet 5: Sonnet costs 7.3x as much
MiniMax M3 costs $0.33 per cached coding session against $2.40 on Claude Sonnet 5. Both have a 1M window, and they differ on tools, caching, and long prompts.
· Prices as of September 28, 2026
MiniMax M3
MiniMax · Released June 1, 2026
MiniMax's open-weight model pitched against closed frontier models, with a 1M context and native multimodality.
MiniMax M3 facts and comparisonsClaude Sonnet 5
Anthropic · Released June 30, 2026
Anthropic's balance of speed and intelligence, and a drop-in upgrade for Claude Sonnet 4.6. Anthropic pitches it as close to Claude Opus 4.8 at a lower price.
Claude Sonnet 5 facts and comparisons
The short answer
MiniMax M3 costs about a seventh as much as Claude Sonnet 5 on the example agentic coding session, $0.33 against $2.40, and the gap stays between 6.7x and 7.8x on every workload. Both accept 1M tokens of context, but MiniMax M3 raises its rates above 512K input tokens and ships open weights, while Sonnet 5 runs in Claude Code, and in Cursor and GitHub Copilot, which don't list MiniMax M3. Choose MiniMax M3 when cost per token decides, and Sonnet 5 when you want Anthropic's model inside those tools.
Choose MiniMax M3 if
- Price decides: $0.30 input and $1.20 output per million tokens, against $2 and $10.
- You want open weights, which MiniMax publishes under the MiniMax community license.
- You need very long outputs: MiniMax lists 524.3K as the maximum and recommends up to 131,072 tokens per request.
- You want image and video input or desktop computer use, which MiniMax names among M3's capabilities.
Choose Claude Sonnet 5 if
- You work in Claude Code, or in Cursor or GitHub Copilot, which don't list MiniMax M3.
- You want one flat rate across the whole 1M window; the price data lists no long-context tier for Sonnet 5.
- You want Anthropic's pitch: performance close to Claude Opus 4.8 at lower prices, with planning and tool use on its own.
- You are still on Claude Sonnet 4.6 and want the drop-in successor Anthropic points to.
Side by side
Specs and prices
| Fact | MiniMax M3 | Claude Sonnet 5 |
|---|---|---|
| Maker | MiniMax | Anthropic |
| API model id | MiniMax-M3 | claude-sonnet-5 |
| Released | June 1, 2026 | June 30, 2026 |
| Status | Current | Current |
| Context window | 1M tokens | 1M tokens |
| Max output | 524.3K tokens | 128K tokens |
| Open weights | Yes | No |
| Input, per 1M tokens | $0.30 | $2 |
| Cache hit, per 1M | $0.06 | $0.20 |
| Cache write, per 1M | $0.30 (same as input) | $2.50 (5-minute), $4 (1-hour) |
| Output, per 1M tokens | $1.20 | $10 |
| Runs in | OpenCode and OpenRouter | Claude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker (MiniMax M3: September 28, 2026; Claude Sonnet 5: September 26, 2026). Batch and priority tiers, taxes, and subscription plans are not included. MiniMax M3: MiniMax labels these rates a permanent 50% discount. Requests over 512K input tokens cost $0.60 input, $0.12 cached, and $2.40 output per million tokens. Claude Sonnet 5: The full 1M context window is billed at standard rates. The launch price became the standard price on August 10, 2026, and a planned increase was cancelled.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | MiniMax M3 | Claude Sonnet 5 |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $0.33 | $2.40 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.06 | $0.40 |
| Output-heavy generation, 30K input, 80K output | $0.11 | $0.86 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $36.30 | $264.00 |
| Where the session’s cost goes | ||
| Cache writes | $0.12 | $1.30 |
| Cache reads | $0.12 | $0.40 |
| Uncached input | $0.03 | $0.20 |
| Output | $0.06 | $0.50 |
- caching saves on the session with MiniMax M3 (59%)
- $0.48
- caching saves on the session with Claude Sonnet 5 (56%)
- $3.10
Where a 7.3x session gap comes from
MiniMax M3 lists $0.30 per million input tokens and $1.20 per million output tokens, against $2 and $10 on Claude Sonnet 5. Input costs 6.7x as much on Sonnet 5 and output 8.3x. MiniMax labels its rates a permanent 50% discount.
Cache hits narrow the ratio. A MiniMax M3 cache hit costs $0.06, one-fifth of its input price, while a Sonnet 5 hit costs $0.20, 10% of input. Hits are 3.3x dearer on Sonnet 5, the narrowest gap between the two rate cards. Writes run the other way: MiniMax charges nothing extra to write the cache, so the 400K written tokens cost $0.12, against $1.30 on Sonnet 5 with Anthropic's 5-minute and 1-hour writes.
The session totals $0.33 against $2.40, 7.3x, and a month of 110 sessions costs $36.30 against $264.00. Uncached work sits close to the headline ratios: 6.7x on the large one-off review, $0.06 against $0.40, and 7.8x on the output-heavy generation, $0.11 against $0.86.
Same 1M window, different long-prompt rules
Both models accept 1M tokens of context. MiniMax M3 requests over 512K input tokens cost $0.60 input, $0.12 cached, and $2.40 output per million, twice its standard rates. Even at that tier it stays well below Sonnet 5's standard rates, and the price data here lists no long-context tier for Sonnet 5.
Output limits differ more on paper than in practice. On Sonnet 5 a single response tops out at 128K tokens. MiniMax M3's maximum is 524.3K, but MiniMax recommends up to 131,072 output tokens per request, which lands close to Sonnet 5's limit. MiniMax caches automatically once a request has 512 or more input tokens, and Claude Code manages caching for you on Sonnet 5.
Open weights and where each model runs
MiniMax publishes M3's weights under the MiniMax community license. Open weights let you run the model on hardware you choose, and this post does not estimate what that costs. On OpenRouter, a request for an open-weight model goes to one of several providers, whose prices can differ from MiniMax's own API, and the table uses MiniMax's own price.
MiniMax M3 runs in OpenCode and OpenRouter. Sonnet 5 runs in Claude Code, Anthropic's own coding agent, and in Cursor, OpenCode, OpenRouter, and GitHub Copilot. In Claude Code the sonnet alias resolves to Sonnet 5 on the Anthropic API.
EveryToken prices Sonnet 5 at Anthropic's rates from Claude Code, Cursor, and OpenCode history, and prices MiniMax M3 when you use it through OpenRouter, from OpenRouter's catalog rather than MiniMax's list price.
How MiniMax and Anthropic position these models
MiniMax pitches M3 against closed frontier models and says "M3 reaches frontier-level performance on specialized tasks such as coding and agentic work." It credits MiniMax Sparse Attention for the 1M context and lists image and video input and operating a desktop computer among its capabilities.
Anthropic calls Sonnet 5 its balance of speed and intelligence, with better reasoning, tool use, and coding than Claude Sonnet 4.6. Adaptive thinking is on by default at high effort. The two makers use different tokenizers, so the same code will not count as the same number of tokens on each, and the 7.3x session gap will not come out exactly 7.3x on your own work.
Prompt caching
How each maker bills cached tokens
MiniMax
MiniMax M3 caches repeated prompts automatically for requests of 512 or more input tokens, with no charge for cache writes.
A cache hit costs $0.06 per million tokens, one-fifth of the input price.
Source: MiniMax docs: Prompt caching
Anthropic
Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.
A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.
Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.
In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.
Source: Anthropic: Prompt caching
Your own numbers
See what MiniMax M3 and Claude Sonnet 5 really cost you.
everyaitoken reads your OpenRouter, Claude Code, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
How much cheaper is MiniMax M3 than Claude Sonnet 5?
On the example agentic coding session, $0.33 against $2.40, or 7.3x. Across the three example workloads the gap runs from 6.7x on a large uncached review to 7.8x on output-heavy generation.
Is MiniMax M3 open source?
MiniMax publishes M3's weights under the MiniMax community license. Claude Sonnet 5 has no open weights. Read the license terms before self-hosting.
Does MiniMax M3 charge more for long prompts?
Yes. Requests over 512K input tokens cost $0.60 input, $0.12 cached, and $2.40 output per million tokens, double its standard rates. Below that, the standard $0.30 input and $1.20 output apply.
Can I use MiniMax M3 in Cursor or GitHub Copilot?
Neither lists it. MiniMax M3 runs in OpenCode and OpenRouter, while Claude Sonnet 5 is in both of those plus Claude Code, Cursor, and GitHub Copilot.
Sources
- MiniMax docs: Pay-as-you-go pricing
- MiniMax: MiniMax M3
- OpenRouter: MiniMax M3
- OpenCode docs: Zen
- Anthropic: Pricing
- Anthropic docs: Claude Sonnet 5
- Anthropic: Introducing Claude Sonnet 5
- Claude Code docs: Model configuration
- Cursor docs: Claude Sonnet 5
- OpenRouter: Claude Sonnet 5
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- MiniMax docs: Prompt caching
- Anthropic: Prompt caching