Model comparison
Qwen3.8-Max vs Claude Opus 5.5: closed models, 2.4x apart
Qwen3.8-Max costs $1.80 on an agentic coding session against $4.40 on Claude Opus 5.5. Both are closed API models, and cache writes drive most of the gap.
· Prices as of September 28, 2026
Qwen3.8-Max
Alibaba Qwen · Released August 2, 2026
Qwen's most capable model, built for long autonomous coding and professional work.
Qwen3.8-Max facts and comparisonsClaude Opus 5.5
Anthropic · Released September 22, 2026
Anthropic's recommended starting model for most work, built for long-running agentic coding. Anthropic says it matches Claude Fable 5.1 on most work at a much lower price.
Claude Opus 5.5 facts and comparisons
The short answer
Qwen3.8-Max costs $1.80 on the example agentic coding session against $4.40 on Claude Opus 5.5, 59% less, with most of the gap coming from Anthropic's cache-write premium and its $20 output rate. Pick Opus 5.5 if you work in Claude Code, Cursor, or GitHub Copilot; pick Qwen3.8-Max for the lower price through OpenRouter or OpenCode, knowing it is a closed API model just as Opus 5.5 is.
Choose Qwen3.8-Max if
- You want the lower price: $2 input and $6 output per million tokens, against $4 and $20.
- You write a lot of new context to the cache, since Qwen's implicit caching bills written tokens as ordinary input.
- You can mark cache breakpoints yourself: Qwen's explicit caching reads at $0.17 per million, below the $0.20 Opus 5.5 charges.
- You work in OpenCode or through OpenRouter, which offer Qwen3.8-Max.
Choose Claude Opus 5.5 if
- You want the model Claude Code starts on by default, which is Opus 5.5 on Pro, Max, Team, and Enterprise plans.
- You pick models in Cursor or GitHub Copilot, which offer Opus 5.5 and do not offer Qwen3.8-Max.
- Speed matters enough to pay for fast mode, Anthropic's research preview for Opus 5.5 at $8 input and $40 output per million.
Side by side
Specs and prices
| Fact | Qwen3.8-Max | Claude Opus 5.5 |
|---|---|---|
| Maker | Alibaba Qwen | Anthropic |
| API model id | qwen3.8-max | claude-opus-5-5 |
| Released | August 2, 2026 | September 22, 2026 |
| Status | Current | Current |
| Context window | 1M tokens | 1M tokens |
| Max output | 131K tokens | 128K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $2 | $4 |
| Cache hit, per 1M | $0.25 | $0.20 |
| Cache write, per 1M | $2 (same as input) | $5 (5-minute), $8 (1-hour) |
| Output, per 1M tokens | $6 | $20 |
| Runs in | OpenCode and OpenRouter | Claude Code, Cursor, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker (Qwen3.8-Max: September 28, 2026; Claude Opus 5.5: September 26, 2026). Batch and priority tiers, taxes, and subscription plans are not included. Qwen3.8-Max: Qwen Cloud prices. Alibaba Cloud Model Studio's international scope lists the same input and output rates. Claude Opus 5.5: The full 1M context window is billed at standard rates. Fast mode, a research preview on the Claude API, costs $8 input and $40 output per million tokens.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Qwen3.8-Max | Claude Opus 5.5 |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $1.80 | $4.40 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.36 | $0.80 |
| Output-heavy generation, 30K input, 80K output | $0.54 | $1.72 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $198.00 | $484.00 |
| Where the session’s cost goes | ||
| Cache writes | $0.80 | $2.60 |
| Cache reads | $0.50 | $0.40 |
| Uncached input | $0.20 | $0.40 |
| Output | $0.30 | $1.00 |
- caching saves on the session with Qwen3.8-Max (66%)
- $3.50
- caching saves on the session with Claude Opus 5.5 (60%)
- $6.60
Qwen3.8-Max is a closed model, like Claude Opus 5.5
Qwen3.8-Max is a closed API model. Qwen publishes open weights for its base variant, Qwen3.8-2.4T-A95B, but not for the Max model you call through the API. Claude Opus 5.5 has no open weights either, so this is a comparison of two hosted services, and self-hosting the exact model is not an option on either side.
Qwen calls Qwen3.8-Max "the most capable model in the Qwen family to date" and pitches it for long autonomous coding and professional work. It says it built a self-evolving harness during an autonomous coding run of more than 10 days, and trained it with reinforcement learning across harnesses including Claude Code and Codex. Alibaba Cloud's Model Studio names Qwen3.8-Max its pick for the strongest reasoning in coding tools.
Anthropic calls Opus 5.5 its recommended starting model for most work and says it performs at the level of Claude Fable 5.1 on most work while costing 40% less to run than Opus 5. Its named strength is long, sprawling work such as codebase-wide migrations and audits. Both makers, in other words, pitch their models at long unattended coding runs.
Where the $2.60 session gap comes from
Qwen3.8-Max charges $2 per million input tokens and $6 per million output tokens. Opus 5.5 charges $4 and $20, which is 2x on input and 3.3x on output. The large one-off review, mostly input, costs $0.36 against $0.80, and the output-heavy generation costs $0.54 against $1.72.
On the agentic session, cache writes dominate. Qwen's implicit cache bills written tokens as ordinary input, so the session's 400K writes cost $0.80. Opus 5.5 writes cost $5 per million for the 5-minute lifetime and $8 for the 1-hour one, 1.25x and 2x its input, which brings the same writes to $2.60. That $1.80 difference is most of the $2.60 gap, and output adds another $0.70.
Cache reads lean slightly toward Opus 5.5. Anthropic bills its hits at 0.05x input, $0.20 per million, while a Qwen implicit hit costs $0.25, or 12.5% of input. Reads cost $0.40 against $0.50 on the session. The totals are $4.40 and $1.80, or $484.00 and $198.00 over 110 sessions a month.
Qwen's two caching modes against Anthropic's
Qwen's implicit caching is automatic and cannot be turned off. It is what this comparison prices: $0.25 per million for a hit, with written tokens at ordinary input. Its explicit caching for Qwen3.8-Max, marked with cache_control, lasts 5 minutes, costs $2.50 per million to create, and costs $0.17 per million to read.
Explicit mode changes the read comparison. At $0.17, an explicit Qwen hit costs less than the $0.20 Opus 5.5 charges, and the $2.50 creation price is half of Anthropic's $5 for a 5-minute write on Opus 5.5. Whether you get those rates depends on whether your tool marks cache breakpoints in its requests.
On the example session, caching saves 66% on Qwen3.8-Max and 60% on Opus 5.5 against sending every token uncached. Anthropic's guidance is that the 5-minute write on Opus 5.5 earns back its premium after a single cache read, and the 1-hour write after two. The homepage cache example walks through the same session.
Where you can use each model
Opus 5.5 is the default model in Claude Code on Pro, Max, Team, and Enterprise plans, and Cursor, OpenCode, OpenRouter, and GitHub Copilot offer it too. Qwen3.8-Max is offered through OpenRouter and in OpenCode. Training across Claude Code and Codex harnesses describes how Qwen trained the model, not where it is offered: both agents are built around their makers' models.
The Qwen prices here are Qwen Cloud's, and Alibaba Cloud Model Studio's international scope lists the same input and output rates. OpenRouter can route a request to a provider whose price differs. EveryToken prices Opus 5.5 at Anthropic's rates from Claude Code, Cursor, and OpenCode, and prices Qwen3.8-Max from OpenRouter's catalog when you use it through OpenRouter.
Both models take 1M tokens of context. Qwen3.8-Max writes up to 131K tokens per response and Opus 5.5 up to 128K. Tokenizers differ, and Anthropic notes that its newer tokenizer counts about 30% more tokens than earlier Claude models for the same text, so compare Qwen3.8-Max and Opus 5.5 on real tasks rather than on token prices alone.
Prompt caching
How each maker bills cached tokens
Alibaba Qwen
Implicit caching is automatic and can't be turned off. A hit on Qwen3.8-Max costs $0.25 per million tokens, and written tokens cost ordinary input.
Explicit caching, marked with cache_control, lasts 5 minutes. On Qwen3.8-Max it costs $2.50 per million tokens to create and $0.17 per million to read. This blog prices Qwen with implicit caching.
Anthropic
Claude caches a prompt prefix up to a breakpoint. One top-level cache_control field places the breakpoint automatically and moves it as the conversation grows, or you can mark up to 4 blocks yourself. Claude Code manages caching for you.
A 5-minute cache write costs 1.25x the input price and a 1-hour write costs 2x. A cache hit costs 0.1x input on most models, 0.05x on Claude Opus 5.5, and 0.025x on Claude Fable 5.1 and Claude Mythos 5.1. Every hit restarts the cache lifetime at no charge.
Anthropic's rule of thumb: a 5-minute write pays for itself after one cache read, and a 1-hour write after two.
In Claude Code, the main conversation uses the 1-hour cache on a Claude subscription and the 5-minute cache with an API key. Each model has its own cache, so switching models starts over.
Source: Anthropic: Prompt caching
Your own numbers
See what Qwen3.8-Max and Claude Opus 5.5 really cost you.
everyaitoken reads your OpenRouter, Claude Code, Cursor, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Does Qwen3.8-Max have open weights?
No. The API model is closed. Qwen publishes open weights for its base variant, Qwen3.8-2.4T-A95B, not for the Max model served through Qwen Cloud or OpenRouter.
Is Qwen3.8-Max cheaper than Claude Opus 5.5?
Yes. The example agentic session costs $1.80 against $4.40, the uncached review $0.36 against $0.80, and output-heavy generation $0.54 against $1.72. Opus 5.5 charges less for an automatic cache hit, $0.20 against $0.25 per million.
Is Qwen3.8-Max in Claude Code's model list?
The sources behind this page don't say. Claude Code is built around Anthropic's models, with Opus 5.5 as its default. Qwen says it trained Qwen3.8-Max across harnesses including Claude Code, and the model is offered through OpenRouter, in OpenCode, and on Qwen's own API.
What does Qwen's explicit caching cost?
On Qwen3.8-Max, explicit caching lasts 5 minutes and costs $2.50 per million tokens to create and $0.17 per million to read. This comparison prices Qwen's implicit caching instead, where a hit costs $0.25 and writes cost ordinary input.
Sources
- Qwen Cloud: Qwen3.8-Max
- Qwen: Qwen3.8
- Alibaba Cloud Model Studio: Model pricing
- OpenRouter: Qwen3.8-Max
- OpenCode docs: Zen
- Anthropic: Pricing
- Anthropic docs: Claude Opus 5.5
- Anthropic: Claude Opus 5.5
- Anthropic docs: Fast mode
- Claude Code docs: Model configuration
- Cursor docs: Claude Opus 5.5
- OpenRouter: Claude Opus 5.5
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Alibaba Cloud Model Studio: Context cache
- Anthropic: Prompt caching