Skip to content

Model comparison

Mistral Medium 3.5 vs Claude Sonnet 5: the cache-write gap

Mistral Medium 3.5 lists rates 25% below Claude Sonnet 5, yet a cached coding session costs 40% less, $1.43 against $2.40. Cache writes explain the difference.

· Prices as of September 28, 2026

  • Mistral Medium 3.5

    Mistral AI · Released May 22, 2026 · Preview

    Mistral's coding and agent flagship, a dense open-weight model that replaced Devstral 2 in Mistral's Vibe coding agents.

    Mistral Medium 3.5 facts and comparisons
  • Claude Sonnet 5

    Anthropic · Released June 30, 2026

    Anthropic's balance of speed and intelligence, and a drop-in upgrade for Claude Sonnet 4.6. Anthropic pitches it as close to Claude Opus 4.8 at a lower price.

    Claude Sonnet 5 facts and comparisons

The short answer

Mistral Medium 3.5 lists its input, output, and cache-hit rates 25% below Claude Sonnet 5, and the example agentic coding session costs $1.43 against $2.40, 40% less, because Mistral charges no cache-write fee while Anthropic bills writes at 1.25x and 2x input. Sonnet 5 is a current model with a 1M context window that runs in Claude Code, Cursor, and GitHub Copilot. Mistral Medium 3.5 is a public preview with a 256K window and open weights, reachable here through OpenRouter, and it powers Mistral's own Vibe coding agents.

Choose Mistral Medium 3.5 if

  • Lower rates matter: $1.50 input and $7.50 output per million tokens, a quarter below Sonnet 5 on both.
  • Your coding agent writes to the cache constantly, and you would rather those writes cost plain input than a premium.
  • You want downloadable weights: Mistral releases Medium 3.5 under a modified MIT license and says the dense model fits on as few as four GPUs.
  • You already use Mistral's Vibe coding agents, where Medium 3.5 took over from Devstral 2.

Choose Claude Sonnet 5 if

  • You would rather not build on a public preview, and Sonnet 5 has been a current release since June 30, 2026.
  • Your repositories or logs need more than 256K tokens in one request, which Sonnet 5's 1M window allows.
  • You work in Claude Code, or in Cursor, OpenCode, or GitHub Copilot, which carry Sonnet 5 but not Mistral Medium 3.5.
  • You want a published output ceiling of 128K tokens, a figure Mistral does not give for Medium 3.5.

Side by side

Specs and prices

FactMistral Medium 3.5Claude Sonnet 5
MakerMistral AIAnthropic
API model idmistral-medium-3.5claude-sonnet-5
ReleasedMay 22, 2026June 30, 2026
StatusPreviewCurrent
Context window256K tokens1M tokens
Max outputNot published128K tokens
Open weightsYesNo
Input, per 1M tokens$1.50$2
Cache hit, per 1M$0.15$0.20
Cache write, per 1M$1.50 (same as input)$2.50 (5-minute), $4 (1-hour)
Output, per 1M tokens$7.50$10
Runs inOpenRouterClaude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker (Mistral Medium 3.5: September 28, 2026; Claude Sonnet 5: September 26, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Mistral Medium 3.5: Mistral lists no per-model cache price. Its caching docs bill cached tokens at 10% of input, which is the $0.15 shown. Claude Sonnet 5: The full 1M context window is billed at standard rates. The launch price became the standard price on August 10, 2026, and a planned increase was cancelled.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadMistral Medium 3.5Claude Sonnet 5
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$1.43$2.40
Large one-off review, 150K input with no cache hits, 10K output$0.30$0.40
Output-heavy generation, 30K input, 80K output$0.65$0.86
A month of sessions, 110 sessions: 5 a day, 22 working days$156.75$264.00
Where the session’s cost goes
Cache writes$0.60$1.30
Cache reads$0.30$0.40
Uncached input$0.15$0.20
Output$0.38$0.50
caching saves on the session with Mistral Medium 3.5 (65%)
$2.70
caching saves on the session with Claude Sonnet 5 (56%)
$3.10

Why the session gap is wider than the rate gap

Mistral Medium 3.5 charges $1.50 per million input tokens and $7.50 per million output tokens, and Claude Sonnet 5 charges $2 and $10. Mistral lists no per-model cache price, but its caching docs bill cached tokens at 10% of input, which works out to $0.15 against Sonnet 5's $0.20. So input, output, and cache hits are each 25% cheaper on Mistral, and uncached work shows exactly that ratio: the large one-off review costs $0.30 against $0.40.

Cache writes break the pattern. Mistral lists no fee for writing the cache, so the 400K tokens the example session writes cost ordinary input, $0.60. Anthropic prices a 5-minute write at 1.25x input and a 1-hour write at 2x, $2.50 and $4 per million on Sonnet 5, and the same writes come to $1.30 there. That $0.70 accounts for most of the $0.97 between the two sessions.

The session lands at $1.43 on Mistral against $2.40 on Sonnet 5, and 110 sessions a month at $156.75 against $264.00. Caching trims 65% off Mistral's uncached session cost and 56% off Sonnet 5's, and the write premium is again why the Anthropic figure is lower.

Mistral Medium 3.5 is a public preview with a 256K window

Mistral announced Medium 3.5 as a public preview, with a release date of May 22, 2026. It took Devstral 2's place in Mistral's Vibe coding agents, and Devstral 2 was retired on July 30, 2026. Mistral's own description is "our frontier-class multimodal model optimized for agentic and coding use cases," and it cites long-horizon tasks and reliable multi-tool calling as strengths.

Its context window is 256K tokens, about a quarter of Sonnet 5's 1M, and Mistral publishes no maximum output. Neither model's price data carries a long-context tier, so each rate holds across its whole window. The limit on Mistral Medium 3.5 is how much fits, not what it costs.

Anthropic positions Sonnet 5 as its balance of speed and intelligence and claims performance close to Claude Opus 4.8 at lower prices. Its launch price became the standard price on August 10, 2026, and Anthropic cancelled a planned increase.

Open weights and the tools that carry each model

Mistral releases Medium 3.5's weights under a modified MIT license and says the dense model can be self-hosted on as few as four GPUs. This post does not estimate what running it yourself would cost. Sonnet 5 has no open weights.

Among the tools compared here, Mistral Medium 3.5 is reachable through OpenRouter. OpenRouter hands a request for an open-weight model to one of several providers, whose prices can differ from Mistral's own API, and the tables use Mistral's price. Sonnet 5 is in Claude Code, Anthropic's own coding agent, and in Cursor, OpenCode, OpenRouter, and GitHub Copilot.

Mistral caches prompt prefixes in 64-token blocks, and passing a prompt_cache_key raises the chance of a hit. EveryToken prices Mistral Medium 3.5 when you use it through OpenRouter, from OpenRouter's catalog, and Sonnet 5 at Anthropic's rates from Claude Code, Cursor, and OpenCode history.

Prompt caching

How each maker bills cached tokens

Mistral AI

Mistral caches prompt prefixes in 64-token blocks. A prompt_cache_key raises the chance of a hit but doesn't guarantee one.

Cached prompt tokens are billed at 10% of the standard input price, and Mistral lists no fee for writing the cache.

Source: Mistral docs: Prompt caching

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

Your own numbers

See what Mistral Medium 3.5 and Claude Sonnet 5 really cost you.

everyaitoken reads your OpenRouter, Claude Code, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is Mistral Medium 3.5 cheaper than Claude Sonnet 5?

Yes. Input, output, and cache-hit rates are 25% lower, and the example agentic session costs $1.43 against $2.40, 40% less, because Mistral adds no cache-write premium. A month of 110 sessions comes to $156.75 against $264.00.

Is Mistral Medium 3.5 still in preview?

Mistral announced it as a public preview. It has nonetheless replaced Devstral 2, retired on July 30, 2026, in Mistral's Vibe coding agents.

How does Mistral price cache hits on Medium 3.5?

Mistral lists no per-model cache price. Its caching docs bill cached prompt tokens at 10% of the standard input price, $0.15 per million on Mistral Medium 3.5, and list no fee for writing the cache.

How much context does each model accept?

Mistral Medium 3.5 accepts 256K tokens and Claude Sonnet 5 accepts 1M. Neither has a long-context price tier in the data here.

  • Claude Fable 5.1 vs Claude Sonnet 5

    Claude Fable 5.1 lists at 5x the rates of Claude Sonnet 5, but cache hits cost $0.25 against $0.20. Why sessions cost 4.4x, and who Anthropic aims Fable 5.1 at.

  • Claude Opus 5.5 vs Claude Sonnet 5

    Claude Opus 5.5 costs twice as much per token as Claude Sonnet 5, yet cache hits cost $0.20 on both. What that means for agentic coding sessions.

  • Claude Sonnet 5 vs Claude Haiku 4.5

    Claude Haiku 4.5 costs half as much as Claude Sonnet 5, with a 200K context window and a retirement date ahead. When the cheaper model fits a coding workflow.

  • Claude Sonnet 5 vs Claude Sonnet 4.6

    Claude Sonnet 5 costs 33% less per token than Claude Sonnet 4.6, but its tokenizer counts up to 1.35x as many tokens. What the upgrade saves in practice.

  • GLM-5.3-Flash vs Claude Sonnet 5

    GLM-5.3-Flash costs $0.16 for a cached coding session that costs $2.40 on Claude Sonnet 5. What drives a 15x gap, and what price alone can't tell you.

  • Claude Sonnet 5 vs Gemini 3.1 Pro Preview

    Claude Sonnet 5 and Gemini 3.1 Pro Preview both charge $2 input, but Gemini's rates rise past 200K tokens and it is still a preview. Session: $2.40 vs $2.00.