Skip to content

Mistral AI

Mistral Medium 3.5: price, context window, and caching

Mistral's coding and agent flagship, a dense open-weight model that replaced Devstral 2 in Mistral's Vibe coding agents.

Released May 22, 2026 · Prices as of September 28, 2026

In Mistral AI’s words

“Our frontier-class multimodal model optimized for agentic and coding use cases.”

Mistral docs: Mistral Medium 3.5

What Mistral AI says it’s good at

  • Long-horizon tasks and reliable multi-tool calling, powering Vibe remote coding agents Source
  • A dense model that can be self-hosted on as few as four GPUs Source

Facts

Specs and prices

FactMistral Medium 3.5
MakerMistral AI
API model idmistral-medium-3.5
ReleasedMay 22, 2026
StatusPreview
Context window256K tokens
Max outputNot published
Open weightsYes
Input, per 1M tokens$1.50
Cache hit, per 1M$0.15
Cache write, per 1M$1.50 (same as input)
Output, per 1M tokens$7.50
Runs inOpenRouter

Standard API rates in US dollars, as published by the maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Mistral Medium 3.5: Mistral lists no per-model cache price. Its caching docs bill cached tokens at 10% of input, which is the $0.15 shown.

Good to know

  • Open weights under a modified MIT license.
  • Announced as a public preview. Devstral 2, which it replaced, was retired on July 30, 2026.

Cost

What typical work costs

Example token counts at Mistral Medium 3.5’s published rates. On the agentic session, caching saves $2.70 against billing every token as ordinary input.

Example workload costs for Mistral Medium 3.5
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$1.43
Large one-off review, 150K input with no cache hits, 10K output$0.30
Output-heavy generation, 30K input, 80K output$0.65
A month of sessions, 110 sessions: 5 a day, 22 working days$156.75

Prompt caching

How Mistral AI bills cached tokens

Mistral AI

Mistral caches prompt prefixes in 64-token blocks. A prompt_cache_key raises the chance of a hit but doesn't guarantee one.

Cached prompt tokens are billed at 10% of the standard input price, and Mistral lists no fee for writing the cache.

Source: Mistral docs: Prompt caching

Mistral Medium 3.5 compared

  • Mistral Medium 3.5 vs Claude Sonnet 5

    Mistral Medium 3.5 lists rates 25% below Claude Sonnet 5, yet a cached coding session costs 40% less, $1.43 against $2.40. Cache writes explain the difference.

  • Mistral Medium 3.5 vs Gemini 3.8 Flash

    Gemini 3.8 Flash costs half as much as Mistral Medium 3.5 on every rate until December 31, 2026. From January 1, 2027, their list prices match to the cent.

  • Mistral Medium 3.5 vs GPT-6 Sol

    Mistral Medium 3.5 undercuts GPT-6 Sol by 25% per token and 32% on a cached coding session. The trade-offs: a 256K window, preview status, and fewer tools.

Your own numbers

See what Mistral Medium 3.5 really costs you.

everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math