Skip to content

Model comparison

Qwen3.8-Max or Claude Sonnet 5? Same input price, compared

Qwen3.8-Max and Claude Sonnet 5 both charge $2 per million input tokens. Output and cache writes set them apart: $1.80 against $2.40 per coding session.

· Prices as of September 28, 2026

  • Qwen3.8-Max

    Alibaba Qwen · Released August 2, 2026

    Qwen's most capable model, built for long autonomous coding and professional work.

    Qwen3.8-Max facts and comparisons
  • Claude Sonnet 5

    Anthropic · Released June 30, 2026

    Anthropic's balance of speed and intelligence, and a drop-in upgrade for Claude Sonnet 4.6. Anthropic pitches it as close to Claude Opus 4.8 at a lower price.

    Claude Sonnet 5 facts and comparisons

The short answer

Qwen3.8-Max and Claude Sonnet 5 charge the same $2 per million input tokens, so the gap comes from output, $6 against $10, and from cache writes, and the example agentic coding session costs $1.80 on Qwen3.8-Max against $2.40 on Sonnet 5. Pick Sonnet 5 if you work in Claude Code, Cursor, or GitHub Copilot; pick Qwen3.8-Max for the lower session cost through OpenRouter or OpenCode.

Choose Qwen3.8-Max if

  • You want output at $6 per million tokens instead of $10, which matters most for long generated files.
  • Your agent writes a lot of new context to the cache, which Qwen's implicit caching bills at the $2 input rate.
  • You control cache breakpoints: Qwen's explicit cache costs $2.50 per million to create, the same as a 5-minute write on Sonnet 5, and $0.17 to read.

Choose Claude Sonnet 5 if

  • You want Sonnet 5 inside Claude Code itself, Anthropic's own agent, where the sonnet alias resolves to it.
  • You use Cursor or GitHub Copilot, which offer Sonnet 5 and do not offer Qwen3.8-Max.
  • Your requests are mostly uncached reading with little output, where price barely separates them: the large review costs $0.40 against $0.36.
  • You reread cached context far more than the example does and want the lower automatic cache-hit price, $0.20 per million against $0.25.

Side by side

Specs and prices

FactQwen3.8-MaxClaude Sonnet 5
MakerAlibaba QwenAnthropic
API model idqwen3.8-maxclaude-sonnet-5
ReleasedAugust 2, 2026June 30, 2026
StatusCurrentCurrent
Context window1M tokens1M tokens
Max output131K tokens128K tokens
Open weightsNoNo
Input, per 1M tokens$2$2
Cache hit, per 1M$0.25$0.20
Cache write, per 1M$2 (same as input)$2.50 (5-minute), $4 (1-hour)
Output, per 1M tokens$6$10
Runs inOpenCode and OpenRouterClaude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot

Standard API rates in US dollars, as published by each maker (Qwen3.8-Max: September 28, 2026; Claude Sonnet 5: September 26, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Qwen3.8-Max: Qwen Cloud prices. Alibaba Cloud Model Studio's international scope lists the same input and output rates. Claude Sonnet 5: The full 1M context window is billed at standard rates. The launch price became the standard price on August 10, 2026, and a planned increase was cancelled.

Cost

What the same work costs

The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.

Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.

Example workload costs
WorkloadQwen3.8-MaxClaude Sonnet 5
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$1.80$2.40
Large one-off review, 150K input with no cache hits, 10K output$0.36$0.40
Output-heavy generation, 30K input, 80K output$0.54$0.86
A month of sessions, 110 sessions: 5 a day, 22 working days$198.00$264.00
Where the session’s cost goes
Cache writes$0.80$1.30
Cache reads$0.50$0.40
Uncached input$0.20$0.20
Output$0.30$0.50
caching saves on the session with Qwen3.8-Max (66%)
$3.50
caching saves on the session with Claude Sonnet 5 (56%)
$3.10

Same input price, different output and cache prices

Qwen3.8-Max and Claude Sonnet 5 both charge $2 per million input tokens. Everything else differs. Output is $6 per million on Qwen3.8-Max and $10 on Sonnet 5, 40% less on Qwen. A Qwen implicit cache hit costs $0.25 and a Sonnet 5 hit $0.20, 20% less on Sonnet. Qwen bills written cache tokens as input, while Anthropic charges $2.50 for a 5-minute write and $4 for a 1-hour write.

Work that is mostly input shows almost no gap. The large one-off review, 150K tokens in and 10K out, costs $0.36 on Qwen3.8-Max and $0.40 on Sonnet 5, and all of the $0.04 difference is output. The output-heavy generation, 30K in and 80K out, costs $0.54 against $0.86.

What decides the cost of an agentic session

The example session costs $1.80 on Qwen3.8-Max and $2.40 on Sonnet 5. Fresh input is identical at $0.20. Output adds $0.20 on Sonnet 5, and cache reads take back $0.10, at $0.40 against $0.50 on Qwen. That leaves the cache writes, $0.80 against $1.30, as the source of most of the $0.60 difference.

Those writes reflect Anthropic's two lifetimes. For Sonnet 5, the session's 400K written tokens are split, half at the 5-minute rate of 1.25x input and half at the 1-hour rate of 2x. Claude Code writes the main conversation to the 1-hour cache on a Claude subscription and to the 5-minute cache with an API key, so how you sign in sets the split. With only 5-minute writes, Sonnet 5 would sit closer to Qwen3.8-Max.

Over 110 sessions a month the totals are $198.00 against $264.00, a $66.00 gap. Caching saves 66% on Qwen3.8-Max and 56% on Sonnet 5 against sending every token uncached. These are API-equivalent estimates at published rates, and the homepage cache example shows how the Qwen3.8-Max and Sonnet 5 session adds up.

Qwen's explicit cache mirrors Anthropic's 5-minute pricing

Qwen offers a second caching mode. Explicit caching, marked with cache_control as on Claude, lasts 5 minutes and costs $2.50 per million to create and $0.17 per million to read. On Qwen3.8-Max that creation price is 1.25x input, the multiple Anthropic uses for its 5-minute write, and it matches Sonnet 5's 5-minute write exactly.

The read is where the modes part ways. An explicit Qwen hit at $0.17 costs less than Qwen's implicit hit at $0.25 and less than Sonnet 5's $0.20. This comparison prices Qwen's implicit mode, which is automatic and cannot be turned off, so a tool that marks explicit breakpoints could pay less on reads than shown here.

By Anthropic's rule of thumb, a 5-minute write on Sonnet 5 pays for itself after one cache read, and each hit restarts its lifetime for free. On Qwen3.8-Max and Sonnet 5 alike, the savings depend on how often your tool resends the same prefix.

Access, limits, and what each maker claims

Anthropic's model reaches further: Claude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot all offer Sonnet 5, while Qwen3.8-Max is on OpenRouter and in OpenCode. Qwen's reinforcement learning across the Claude Code and Codex harnesses is a training claim, not a statement about either tool's model list.

Neither model has open weights. Qwen publishes open weights for the base variant, Qwen3.8-2.4T-A95B, and keeps the Max API model closed. Both take 1M tokens of context, and Qwen3.8-Max writes up to 131K tokens per response against 128K.

Anthropic says "Claude Sonnet 5 is built to be the most agentic Sonnet model yet" and pitches it as close to Claude Opus 4.8 at lower prices, with adaptive thinking on by default at high effort. Qwen calls Qwen3.8-Max the most capable model in its family and points to an autonomous coding run of more than 10 days. EveryToken prices Sonnet 5 from Claude Code, Cursor, and OpenCode history at Anthropic's rates, and Qwen3.8-Max from OpenRouter's catalog when you use it through OpenRouter.

Prompt caching

How each maker bills cached tokens

Alibaba Qwen

Implicit caching is automatic and can't be turned off. A hit on Qwen3.8-Max costs $0.25 per million tokens, and written tokens cost ordinary input.

Explicit caching, marked with cache_control, lasts 5 minutes. On Qwen3.8-Max it costs $2.50 per million tokens to create and $0.17 per million to read. This blog prices Qwen with implicit caching.

Source: Alibaba Cloud Model Studio: Context cache

Anthropic

Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.

A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.

Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.

In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.

Source: Anthropic: Prompt caching

Your own numbers

See what Qwen3.8-Max and Claude Sonnet 5 really cost you.

everyaitoken reads your OpenRouter, Claude Code, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math

FAQ

Questions

Do Qwen3.8-Max and Claude Sonnet 5 cost the same?

They share the $2 input price, but Qwen3.8-Max is cheaper on output, $6 against $10, and on cache writes. The example agentic session costs $1.80 on Qwen3.8-Max and $2.40 on Sonnet 5.

Which model has cheaper cache hits?

With automatic caching, Sonnet 5, at $0.20 per million against $0.25. With Qwen's explicit caching, Qwen3.8-Max, at $0.17 per million.

Is Qwen3.8-Max an open-weight model?

No. The API model is closed. Qwen publishes open weights for its base variant, Qwen3.8-2.4T-A95B, not for the Max model you call.

Where can I use Qwen3.8-Max?

Through OpenRouter, in OpenCode, and on Qwen Cloud, whose rates this page uses. Alibaba Cloud Model Studio lists the same Qwen3.8-Max input and output rates in its international scope. Cursor and GitHub Copilot do not offer Qwen3.8-Max, and Claude Code is built around Anthropic's models.

  • Claude Fable 5.1 vs Claude Sonnet 5

    Claude Fable 5.1 lists at 5x the rates of Claude Sonnet 5, but cache hits cost $0.25 against $0.20. Why sessions cost 4.4x, and who Anthropic aims Fable 5.1 at.

  • Claude Opus 5.5 vs Claude Sonnet 5

    Claude Opus 5.5 costs twice as much per token as Claude Sonnet 5, yet cache hits cost $0.20 on both. What that means for agentic coding sessions.

  • Claude Sonnet 5 vs Claude Haiku 4.5

    Claude Haiku 4.5 costs half as much as Claude Sonnet 5, with a 200K context window and a retirement date ahead. When the cheaper model fits a coding workflow.

  • Claude Sonnet 5 vs Claude Sonnet 4.6

    Claude Sonnet 5 costs 33% less per token than Claude Sonnet 4.6, but its tokenizer counts up to 1.35x as many tokens. What the upgrade saves in practice.

  • Qwen3.8-Max vs Claude Opus 5.5

    Qwen3.8-Max costs $1.80 on an agentic coding session against $4.40 on Claude Opus 5.5. Both are closed API models, and cache writes drive most of the gap.

  • GLM-5.3-Flash vs Claude Sonnet 5

    GLM-5.3-Flash costs $0.16 for a cached coding session that costs $2.40 on Claude Sonnet 5. What drives a 15x gap, and what price alone can't tell you.