Skip to content

Alibaba Qwen

Qwen3.8-Max: price, context window, and caching

Qwen's most capable model, built for long autonomous coding and professional work.

Released August 2, 2026 · Prices as of September 28, 2026

In Alibaba Qwen’s words

“Today, we are officially releasing Qwen 3.8-Max, the most capable model in the Qwen family to date.”

Qwen: Qwen3.8

What Alibaba Qwen says it’s good at

  • A self-evolving harness built during an autonomous coding run of more than 10 days Source
  • Reinforcement learning across harnesses including Claude Code and Codex Source
  • Model Studio's pick for the strongest reasoning in coding tools Source

Facts

Specs and prices

FactQwen3.8-Max
MakerAlibaba Qwen
API model idqwen3.8-max
ReleasedAugust 2, 2026
StatusCurrent
Context window1M tokens
Max output131K tokens
Open weightsNo
Input, per 1M tokens$2
Cache hit, per 1M$0.25
Cache write, per 1M$2 (same as input)
Output, per 1M tokens$6
Runs inOpenCode and OpenRouter

Standard API rates in US dollars, as published by the maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Qwen3.8-Max: Qwen Cloud prices. Alibaba Cloud Model Studio's international scope lists the same input and output rates.

Good to know

  • The API model is closed, and Qwen publishes open weights for its base variant, Qwen3.8-2.4T-A95B.

Cost

What typical work costs

Example token counts at Qwen3.8-Max’s published rates. On the agentic session, caching saves $3.50 against billing every token as ordinary input.

Example workload costs for Qwen3.8-Max
Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output$1.80
Large one-off review, 150K input with no cache hits, 10K output$0.36
Output-heavy generation, 30K input, 80K output$0.54
A month of sessions, 110 sessions: 5 a day, 22 working days$198.00

Prompt caching

How Alibaba Qwen bills cached tokens

Alibaba Qwen

Implicit caching is automatic and can't be turned off. A hit on Qwen3.8-Max costs $0.25 per million tokens, and written tokens cost ordinary input.

Explicit caching, marked with cache_control, lasts 5 minutes. On Qwen3.8-Max it costs $2.50 per million tokens to create and $0.17 per million to read. This blog prices Qwen with implicit caching.

Source: Alibaba Cloud Model Studio: Context cache

Qwen3.8-Max compared

  • Qwen3.8-Max vs Claude Sonnet 5

    Qwen3.8-Max and Claude Sonnet 5 both charge $2 per million input tokens. Output and cache writes set them apart: $1.80 against $2.40 per coding session.

  • Qwen3.8-Max vs Claude Opus 5.5

    Qwen3.8-Max costs $1.80 on an agentic coding session against $4.40 on Claude Opus 5.5. Both are closed API models, and cache writes drive most of the gap.

  • Qwen3.8-Max vs GLM-5.3

    GLM-5.3 costs $1.44 per cached coding session against $1.80 on Qwen3.8-Max, even though cache hits cost about the same. Rates, caching, weights, and tools.

  • Qwen3.8-Max vs GPT-6 Sol

    Qwen3.8-Max and GPT-6 Sol list the same $2 input rate, and an agentic coding session differs by $0.30. How output, cache writes, and Codex shape the choice.

  • Qwen3.8-Max vs Kimi K3

    Qwen3.8-Max undercuts Kimi K3 on every rate, $1.80 against $2.85 for an agentic coding session. Kimi K3 has open weights; the Qwen3.8-Max API model does not.

Your own numbers

See what Qwen3.8-Max really costs you.

everyaitoken reads your OpenRouter history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.

Launching soonSee the cache math