Model comparison
Qwen3.8-Max or Claude Sonnet 5? Same input price, compared
Qwen3.8-Max and Claude Sonnet 5 both charge $2 per million input tokens. Output and cache writes set them apart: $1.80 against $2.40 per coding session.
· Prices as of September 28, 2026
Qwen3.8-Max
Alibaba Qwen · Released August 2, 2026
Qwen's most capable model, built for long autonomous coding and professional work.
Qwen3.8-Max facts and comparisonsClaude Sonnet 5
Anthropic · Released June 30, 2026
Anthropic's balance of speed and intelligence, and a drop-in upgrade for Claude Sonnet 4.6. Anthropic pitches it as close to Claude Opus 4.8 at a lower price.
Claude Sonnet 5 facts and comparisons
The short answer
Qwen3.8-Max and Claude Sonnet 5 charge the same $2 per million input tokens, so the gap comes from output, $6 against $10, and from cache writes, and the example agentic coding session costs $1.80 on Qwen3.8-Max against $2.40 on Sonnet 5. Pick Sonnet 5 if you work in Claude Code, Cursor, or GitHub Copilot; pick Qwen3.8-Max for the lower session cost through OpenRouter or OpenCode.
Choose Qwen3.8-Max if
- You want output at $6 per million tokens instead of $10, which matters most for long generated files.
- Your agent writes a lot of new context to the cache, which Qwen's implicit caching bills at the $2 input rate.
- You control cache breakpoints: Qwen's explicit cache costs $2.50 per million to create, the same as a 5-minute write on Sonnet 5, and $0.17 to read.
Choose Claude Sonnet 5 if
- You want Sonnet 5 inside Claude Code itself, Anthropic's own agent, where the sonnet alias resolves to it.
- You use Cursor or GitHub Copilot, which offer Sonnet 5 and do not offer Qwen3.8-Max.
- Your requests are mostly uncached reading with little output, where price barely separates them: the large review costs $0.40 against $0.36.
- You reread cached context far more than the example does and want the lower automatic cache-hit price, $0.20 per million against $0.25.
Side by side
Specs and prices
| Fact | Qwen3.8-Max | Claude Sonnet 5 |
|---|---|---|
| Maker | Alibaba Qwen | Anthropic |
| API model id | qwen3.8-max | claude-sonnet-5 |
| Released | August 2, 2026 | June 30, 2026 |
| Status | Current | Current |
| Context window | 1M tokens | 1M tokens |
| Max output | 131K tokens | 128K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $2 | $2 |
| Cache hit, per 1M | $0.25 | $0.20 |
| Cache write, per 1M | $2 (same as input) | $2.50 (5-minute), $4 (1-hour) |
| Output, per 1M tokens | $6 | $10 |
| Runs in | OpenCode and OpenRouter | Claude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker (Qwen3.8-Max: September 28, 2026; Claude Sonnet 5: September 26, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Qwen3.8-Max: Qwen Cloud prices. Alibaba Cloud Model Studio's international scope lists the same input and output rates. Claude Sonnet 5: The full 1M context window is billed at standard rates. The launch price became the standard price on August 10, 2026, and a planned increase was cancelled.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Qwen3.8-Max | Claude Sonnet 5 |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $1.80 | $2.40 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.36 | $0.40 |
| Output-heavy generation, 30K input, 80K output | $0.54 | $0.86 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $198.00 | $264.00 |
| Where the session’s cost goes | ||
| Cache writes | $0.80 | $1.30 |
| Cache reads | $0.50 | $0.40 |
| Uncached input | $0.20 | $0.20 |
| Output | $0.30 | $0.50 |
- caching saves on the session with Qwen3.8-Max (66%)
- $3.50
- caching saves on the session with Claude Sonnet 5 (56%)
- $3.10
Same input price, different output and cache prices
Qwen3.8-Max and Claude Sonnet 5 both charge $2 per million input tokens. Everything else differs. Output is $6 per million on Qwen3.8-Max and $10 on Sonnet 5, 40% less on Qwen. A Qwen implicit cache hit costs $0.25 and a Sonnet 5 hit $0.20, 20% less on Sonnet. Qwen bills written cache tokens as input, while Anthropic charges $2.50 for a 5-minute write and $4 for a 1-hour write.
Work that is mostly input shows almost no gap. The large one-off review, 150K tokens in and 10K out, costs $0.36 on Qwen3.8-Max and $0.40 on Sonnet 5, and all of the $0.04 difference is output. The output-heavy generation, 30K in and 80K out, costs $0.54 against $0.86.
What decides the cost of an agentic session
The example session costs $1.80 on Qwen3.8-Max and $2.40 on Sonnet 5. Fresh input is identical at $0.20. Output adds $0.20 on Sonnet 5, and cache reads take back $0.10, at $0.40 against $0.50 on Qwen. That leaves the cache writes, $0.80 against $1.30, as the source of most of the $0.60 difference.
Those writes reflect Anthropic's two lifetimes. For Sonnet 5, the session's 400K written tokens are split, half at the 5-minute rate of 1.25x input and half at the 1-hour rate of 2x. Claude Code writes the main conversation to the 1-hour cache on a Claude subscription and to the 5-minute cache with an API key, so how you sign in sets the split. With only 5-minute writes, Sonnet 5 would sit closer to Qwen3.8-Max.
Over 110 sessions a month the totals are $198.00 against $264.00, a $66.00 gap. Caching saves 66% on Qwen3.8-Max and 56% on Sonnet 5 against sending every token uncached. These are API-equivalent estimates at published rates, and the homepage cache example shows how the Qwen3.8-Max and Sonnet 5 session adds up.
Qwen's explicit cache mirrors Anthropic's 5-minute pricing
Qwen offers a second caching mode. Explicit caching, marked with cache_control as on Claude, lasts 5 minutes and costs $2.50 per million to create and $0.17 per million to read. On Qwen3.8-Max that creation price is 1.25x input, the multiple Anthropic uses for its 5-minute write, and it matches Sonnet 5's 5-minute write exactly.
The read is where the modes part ways. An explicit Qwen hit at $0.17 costs less than Qwen's implicit hit at $0.25 and less than Sonnet 5's $0.20. This comparison prices Qwen's implicit mode, which is automatic and cannot be turned off, so a tool that marks explicit breakpoints could pay less on reads than shown here.
By Anthropic's rule of thumb, a 5-minute write on Sonnet 5 pays for itself after one cache read, and each hit restarts its lifetime for free. On Qwen3.8-Max and Sonnet 5 alike, the savings depend on how often your tool resends the same prefix.
Access, limits, and what each maker claims
Anthropic's model reaches further: Claude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot all offer Sonnet 5, while Qwen3.8-Max is on OpenRouter and in OpenCode. Qwen's reinforcement learning across the Claude Code and Codex harnesses is a training claim, not a statement about either tool's model list.
Neither model has open weights. Qwen publishes open weights for the base variant, Qwen3.8-2.4T-A95B, and keeps the Max API model closed. Both take 1M tokens of context, and Qwen3.8-Max writes up to 131K tokens per response against 128K.
Anthropic says "Claude Sonnet 5 is built to be the most agentic Sonnet model yet" and pitches it as close to Claude Opus 4.8 at lower prices, with adaptive thinking on by default at high effort. Qwen calls Qwen3.8-Max the most capable model in its family and points to an autonomous coding run of more than 10 days. EveryToken prices Sonnet 5 from Claude Code, Cursor, and OpenCode history at Anthropic's rates, and Qwen3.8-Max from OpenRouter's catalog when you use it through OpenRouter.
Prompt caching
How each maker bills cached tokens
Alibaba Qwen
Implicit caching is automatic and can't be turned off. A hit on Qwen3.8-Max costs $0.25 per million tokens, and written tokens cost ordinary input.
Explicit caching, marked with cache_control, lasts 5 minutes. On Qwen3.8-Max it costs $2.50 per million tokens to create and $0.17 per million to read. This blog prices Qwen with implicit caching.
Anthropic
Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.
A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.
Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.
In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.
Source: Anthropic: Prompt caching
Your own numbers
See what Qwen3.8-Max and Claude Sonnet 5 really cost you.
everyaitoken reads your OpenRouter, Claude Code, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Do Qwen3.8-Max and Claude Sonnet 5 cost the same?
They share the $2 input price, but Qwen3.8-Max is cheaper on output, $6 against $10, and on cache writes. The example agentic session costs $1.80 on Qwen3.8-Max and $2.40 on Sonnet 5.
Which model has cheaper cache hits?
With automatic caching, Sonnet 5, at $0.20 per million against $0.25. With Qwen's explicit caching, Qwen3.8-Max, at $0.17 per million.
Is Qwen3.8-Max an open-weight model?
No. The API model is closed. Qwen publishes open weights for its base variant, Qwen3.8-2.4T-A95B, not for the Max model you call.
Where can I use Qwen3.8-Max?
Through OpenRouter, in OpenCode, and on Qwen Cloud, whose rates this page uses. Alibaba Cloud Model Studio lists the same Qwen3.8-Max input and output rates in its international scope. Cursor and GitHub Copilot do not offer Qwen3.8-Max, and Claude Code is built around Anthropic's models.
Sources
- Qwen Cloud: Qwen3.8-Max
- Qwen: Qwen3.8
- Alibaba Cloud Model Studio: Model pricing
- OpenRouter: Qwen3.8-Max
- OpenCode docs: Zen
- Anthropic: Pricing
- Anthropic docs: Claude Sonnet 5
- Anthropic: Introducing Claude Sonnet 5
- Claude Code docs: Model configuration
- Cursor docs: Claude Sonnet 5
- OpenRouter: Claude Sonnet 5
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Alibaba Cloud Model Studio: Context cache
- Anthropic: Prompt caching