Skip to content

Model comparison

Mistral Medium 3.5 vs Gemini 3.8 Flash: the 2027 price match

Gemini 3.8 Flash costs half as much as Mistral Medium 3.5 on every rate until December 31, 2026. From January 1, 2027, their list prices match to the cent.

· Prices as of September 28, 2026

  • Mistral Medium 3.5

    Mistral AI · Released May 22, 2026 · Preview

    Mistral's coding and agent flagship, a dense open-weight model that replaced Devstral 2 in Mistral's Vibe coding agents.

    Mistral Medium 3.5 facts and comparisons
  • Gemini 3.8 Flash

    Google · Released September 2, 2026

    Google's newest and most capable Flash model, aimed at long-running coding and agent work at Flash prices.

    Gemini 3.8 Flash facts and comparisons

The short answer

Gemini 3.8 Flash costs exactly half as much as Mistral Medium 3.5 on every rate today, so the example agentic coding session runs $0.71 against $1.43. That changes on January 1, 2027, when Flash's introductory rates end and its new list prices, $1.50 input, $0.15 cached, and $7.50 output, equal Mistral Medium 3.5's rates. From then on the choice rests on limits and tools: Flash has a 1.05M window and runs in Gemini CLI, Cursor, and GitHub Copilot, while Mistral Medium 3.5 has open weights and a 256K window.

Choose Mistral Medium 3.5 if

  • You want weights you can run yourself, which Mistral publishes under a modified MIT license.
  • You are planning for 2027, when Flash's list price rises to Mistral Medium 3.5's level and the price advantage goes away.
  • You use Mistral's Vibe coding agents, which Medium 3.5 now powers.
  • You are drawn to Mistral's description of a frontier-class multimodal model tuned for agentic and coding work.

Choose Gemini 3.8 Flash if

  • You want the lower price now: through December 31, 2026, every Flash rate is half of Mistral Medium 3.5's.
  • You need more than 256K tokens of context, and Flash's window reaches 1.05M.
  • You work in Gemini CLI, or in Cursor, OpenCode, or GitHub Copilot, which list Flash but not Mistral Medium 3.5.
  • A zero-cost start helps, and Flash's input, output, and caching are covered by the Gemini API free tier.

Side by side

Specs and prices

FactMistral Medium 3.5Gemini 3.8 Flash
MakerMistral AIGoogle
API model idmistral-medium-3.5gemini-3.8-flash
ReleasedMay 22, 2026September 2, 2026
StatusPreviewCurrent
Context window256K tokens1.05M tokens
Max outputNot published65.5K tokens
Open weightsYesNo
Input, per 1M tokens$1.50$0.75
Cache hit, per 1M$0.15$0.075
Cache write, per 1M$1.50 (same as input)$0.75 (same as input)
Output, per 1M tokens$7.50$3.75
Runs inOpenRouterCursor, Gemini CLI, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Mistral Medium 3.5: Mistral lists no per-model cache price. Its caching docs bill cached tokens at 10% of input, which is the $0.15 shown. Gemini 3.8 Flash: Introductory rates through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached, and $7.50 output per million tokens.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadMistral Medium 3.5Gemini 3.8 Flash
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$1.43$0.71
Large one-off review, 150K input with no cache hits, 10K output$0.30$0.15
Output-heavy generation, 30K input, 80K output$0.65$0.32
A month of sessions, 110 sessions: 5 a day, 22 working days$156.75$78.38
Where the session’s cost goes
Cache writes$0.60$0.30
Cache reads$0.30$0.15
Uncached input$0.15$0.08
Output$0.38$0.19
caching saves on the session with Mistral Medium 3.5 (65%)
$2.70
caching saves on the session with Gemini 3.8 Flash (66%)
$1.35

Half the price now, the same price from January 2027

Put the two rate cards side by side and every Mistral number is double. Mistral Medium 3.5 charges $1.50 per million for input, $0.15 for a cache hit, and $7.50 for output, and Gemini 3.8 Flash charges $0.75, $0.075, and $3.75. Neither maker charges a cache-write fee, so writes follow the same 2x ratio.

The workloads follow the same ratio: $0.71 against $1.43 for the agentic session, $0.15 against $0.30 for the large one-off review, $0.32 against $0.65 for the output-heavy generation, and $78.38 against $156.75 for a month of 110 sessions.

Google labels Flash's current prices introductory rates through December 31, 2026. From January 1, 2027 they rise to $1.50 input, $0.15 cached, and $7.50 output, which are Mistral Medium 3.5's rates to the cent. On the published schedules the price gap closes entirely at the new year, and what separates the two afterwards is everything other than price.

What separates them once the price doesn't

Context is the clearest difference. Flash accepts 1.05M tokens and writes up to 65.5K per request. Mistral Medium 3.5 accepts 256K, and Mistral does not publish a maximum output. Neither model's price data lists a long-context tier.

Status differs too. Flash is a current model, released on September 2, 2026, and Google calls it "our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows." Mistral announced Medium 3.5 as a public preview in May 2026 and describes it as "our frontier-class multimodal model optimized for agentic and coding use cases." In Mistral's Vibe coding agents it replaced Devstral 2, which was retired on July 30, 2026.

Then there are the weights. Mistral publishes Medium 3.5's under a modified MIT license and says the dense model runs on as few as four GPUs. Flash's weights are not published, and this post does not estimate hosting costs for either.

Caching rules and where each model runs

Mistral publishes no per-model cache price; the $0.15 in the tables comes from its rule that cached prompt tokens cost 10% of input, the same share Google charges on Flash. Neither maker lists a write fee, so caching saves 65% of Mistral's uncached session and 66% of Flash's. The mechanics differ. Mistral caches prefixes in 64-token blocks, and a prompt_cache_key improves the odds of a hit. Google's implicit caching is on by default without a promised hit, and its explicit caching adds storage at $0.50 to $1 per million tokens per hour on Flash models.

For Gemini API key and Vertex AI users, Gemini CLI, Google's own coding agent, uses Flash for the Flash side of its default auto model. Cursor, OpenCode, OpenRouter, and GitHub Copilot list it as well. Mistral Medium 3.5 is reachable here through OpenRouter, which routes open-weight requests to one of several providers whose prices can differ from Mistral's own API.

EveryToken prices Flash at Google's rates from Gemini CLI, Cursor, or OpenCode history, and prices Mistral Medium 3.5 from OpenRouter's catalog when it runs through OpenRouter, so a session on either shows up as an API-equivalent estimate.

Prompt caching

How each maker bills cached tokens

Mistral AI

Mistral caches prompt prefixes in 64-token blocks. A prompt_cache_key raises the chance of a hit but doesn't guarantee one.

Cached prompt tokens are billed at 10% of the standard input price, and Mistral lists no fee for writing the cache.

Source: Mistral docs: Prompt caching

Google

Implicit caching is on by default for Gemini 2.5 and newer. When a request repeats a prefix Google has cached, the discount is applied automatically, but a hit is not guaranteed.

Explicit caching creates a cache you reference by name, with a guaranteed discount. It adds a storage charge for as long as the cache lives, 1 hour by default: $4.50 per million tokens per hour on Pro models and $0.50 to $1 on Flash models.

On every model compared here, a cache hit costs 10% of the input price. Google publishes no separate cache-write price, so this blog prices written tokens as ordinary input.

Source: Google: Context caching

Your own numbers

See what Mistral Medium 3.5 and Gemini 3.8 Flash really cost you.

everyaitoken reads your OpenRouter, Gemini CLI, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Is Gemini 3.8 Flash half the price of Mistral Medium 3.5?

Through December 31, 2026, yes, on every rate: $0.75 against $1.50 for input, $0.075 against $0.15 for cache hits, and $3.75 against $7.50 for output.

What will Gemini 3.8 Flash cost in 2027?

From January 1, 2027, $1.50 input, $0.15 cached, and $7.50 output per million tokens. Those are the same rates Mistral Medium 3.5 lists today.

Which of the two can read a larger codebase in one request?

Gemini 3.8 Flash, with a 1.05M context window against 256K on Mistral Medium 3.5. Flash caps each response at 65.5K tokens, and Mistral doesn't publish an output cap.

Can I run Mistral Medium 3.5 myself?

Mistral publishes its weights under a modified MIT license and says the dense model can be self-hosted on as few as four GPUs. This post does not estimate what that would cost.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash

    Gemini 3.8 Flash costs about half of Gemini 3.5 Flash on introductory rates. What changes on January 1, 2027, and which one Gemini CLI picks for you.

  • Gemini 3.8 Flash vs Gemini 3.5 Flash-Lite

    Google recommends Gemini 3.8 Flash and Gemini 3.5 Flash-Lite together for new projects. Where each fits, the 2.1x session gap, and what changes in 2027.

  • Gemini 3.8 Flash vs Gemini 3.1 Pro Preview

    Gemini CLI's auto model uses both Gemini 3.8 Flash and Gemini 3.1 Pro Preview. What each costs, why Flash rates double in 2027, and the Pro's preview status.

  • Gemini 3.8 Flash vs Gemini 3.7 Flash

    Gemini 3.8 Flash and Gemini 3.7 Flash cost the same, and both rates double on January 1, 2027. What differs is Gemini CLI, Cursor, and Google's own pitch.

  • Claude Fable 5.1 vs Gemini 3.8 Flash

    A cached coding session costs 14.8x more on Claude Fable 5.1 than on Gemini 3.8 Flash. How Flash's introductory rates, output cap, and free tier factor in.

  • Claude Haiku 4.5 vs Gemini 3.8 Flash

    Gemini 3.8 Flash undercuts Claude Haiku 4.5 on introductory rates, $0.71 against $1.20 per cached session. From January 1, 2027, its rates top Haiku's.