Model comparison
Qwen3.8-Max vs GPT-6 Sol: equal input, $6 or $10 output
Qwen3.8-Max and GPT-6 Sol list the same $2 input rate, and an agentic coding session differs by $0.30. How output, cache writes, and Codex shape the choice.
· Prices as of September 28, 2026
Qwen3.8-Max
Alibaba Qwen · Released August 2, 2026
Qwen's most capable model, built for long autonomous coding and professional work.
Qwen3.8-Max facts and comparisonsGPT-6 Sol
OpenAI · Released September 22, 2026
The mid-priced GPT-6 model, which OpenAI pitches for complex coding and agent workflows and which the Codex docs recommend for complex coding.
GPT-6 Sol facts and comparisons
The short answer
Qwen3.8-Max and GPT-6 Sol both charge $2 per million input tokens, and the example agentic coding session costs $1.80 on Qwen3.8-Max against $2.10 on GPT-6 Sol, a 14% gap driven by output and OpenAI's 1.25x cache-write charge. Pick GPT-6 Sol if you work in Codex, whose docs recommend it for complex coding; pick Qwen3.8-Max for cheaper output through OpenRouter or OpenCode.
Choose Qwen3.8-Max if
- Your work is output-heavy: Qwen3.8-Max charges $6 per million output tokens against $10, and the example generation costs $0.54 against $0.86.
- You want cache writes without a premium, which Qwen's implicit caching bills at the $2 input rate.
- You mark cache breakpoints yourself, where Qwen's explicit hit costs $0.17 per million, below GPT-6 Sol's $0.20.
Choose GPT-6 Sol if
- You code in Codex, which is built around OpenAI's models and names GPT-6 Sol its pick for complex coding.
- You want a cached prefix that stays reusable for at least 30 minutes after its last use, rather than a 5-minute explicit cache.
- You rely on automatic caching and reread far more than the example session does: a GPT-6 Sol hit costs $0.20 per million against $0.25 on Qwen's implicit cache.
- You use GitHub Copilot, which offers GPT-6 Sol and does not offer Qwen3.8-Max.
Side by side
Specs and prices
| Fact | Qwen3.8-Max | GPT-6 Sol |
|---|---|---|
| Maker | Alibaba Qwen | OpenAI |
| API model id | qwen3.8-max | gpt-6-sol |
| Released | August 2, 2026 | September 22, 2026 |
| Status | Current | Current |
| Context window | 1M tokens | 1.05M tokens |
| Max output | 131K tokens | 128K tokens |
| Open weights | No | No |
| Input, per 1M tokens | $2 | $2 |
| Cache hit, per 1M | $0.25 | $0.20 |
| Cache write, per 1M | $2 (same as input) | $2.50 |
| Output, per 1M tokens | $6 | $10 |
| Runs in | OpenCode and OpenRouter | Codex, OpenCode, OpenRouter, and GitHub Copilot |
Standard API rates in US dollars, as published by each maker on September 28, 2026. Batch and priority tiers, taxes, and subscription plans are not included. Qwen3.8-Max: Qwen Cloud prices. Alibaba Cloud Model Studio's international scope lists the same input and output rates. GPT-6 Sol: Requests over 272K input tokens cost 2x for input and cache and 1.5x for output, for the whole request.
Cost
What the same work costs
The same token counts, priced at each model’s published rates. Cached tokens are billed at each maker’s cache prices, so the session shows what caching is worth on each model.
Real sessions differ: the two models count the same code as different numbers of tokens, and reasoning settings change how much each one writes. Your own history is the real test.
| Workload | Qwen3.8-Max | GPT-6 Sol |
|---|---|---|
| Agentic coding session, 100K input, 400K written to cache, 2M read from cache, 50K output | $1.80 | $2.10 |
| Large one-off review, 150K input with no cache hits, 10K output | $0.36 | $0.40 |
| Output-heavy generation, 30K input, 80K output | $0.54 | $0.86 |
| A month of sessions, 110 sessions: 5 a day, 22 working days | $198.00 | $231.00 |
| Where the session’s cost goes | ||
| Cache writes | $0.80 | $1.00 |
| Cache reads | $0.50 | $0.40 |
| Uncached input | $0.20 | $0.20 |
| Output | $0.30 | $0.50 |
- caching saves on the session with Qwen3.8-Max (66%)
- $3.50
- caching saves on the session with GPT-6 Sol (62%)
- $3.40
Where a $0.30 difference per session comes from
Qwen3.8-Max and GPT-6 Sol share a $2 input rate, so fresh input costs the same $0.20 in the example session. Output is the larger split: $6 per million on Qwen3.8-Max against $10 on GPT-6 Sol. On the session's 50K output tokens that is $0.30 against $0.50.
Cache writes add another $0.20. OpenAI bills a write at 1.25x input, $2.50 per million, while Qwen's implicit cache bills written tokens as ordinary input, so the session's 400K writes cost $1.00 against $0.80. Reads go the other way, $0.40 on GPT-6 Sol and $0.50 on Qwen3.8-Max, because OpenAI's $0.20 hit undercuts Qwen's $0.25.
Net, the session costs $2.10 against $1.80, or $231.00 against $198.00 over 110 sessions a month. Uncached work tells the same story: the large one-off review is $0.40 against $0.36, and the output-heavy generation $0.86 against $0.54, 37% less on Qwen3.8-Max.
Qwen's explicit cache against OpenAI's cache
Qwen has two caching modes. The implicit one is automatic, cannot be switched off, and is the one priced against GPT-6 Sol here. The explicit one, marked with cache_control, lasts 5 minutes and charges $2.50 per million to create and $0.17 per million to read.
Set beside OpenAI's rules, the explicit mode looks familiar. Its $2.50 creation price is exactly what GPT-6 Sol charges to write the cache, and its $0.17 read is below GPT-6 Sol's $0.20 hit. The trade is lifetime: OpenAI keeps a cached prefix reusable for at least 30 minutes after its last use, while a Qwen explicit cache lasts 5 minutes.
Caching saves 66% on Qwen3.8-Max and 62% on GPT-6 Sol in the example session, against sending every token uncached. GPT-6 Sol's caching starts at 1,024 input tokens and takes up to four explicit breakpoints. The cache example on our homepage walks through the Qwen3.8-Max and GPT-6 Sol session line by line.
The 272K rule, context, and output limits
GPT-6 Sol has a 1.05M context window and accepts up to 922K input tokens. Beyond 272K input tokens, OpenAI reprices the whole GPT-6 Sol request at 2x for input and cache and 1.5x for output. Qwen3.8-Max has 1M, and the Qwen Cloud rates used here list no long-context tier.
Output limits are close: 131K on Qwen3.8-Max and 128K on GPT-6 Sol. The example session stays under 200K per request, so the 272K line does not touch it, but an agent that loads a whole repository into one prompt can cross it on GPT-6 Sol.
GPT-6 Sol defaults to medium reasoning effort in the API and in Codex. Output is 24% of its session cost here and 17% of the Qwen3.8-Max session, so reasoning length is a real lever on both.
Codex, OpenCode, and how OpenAI and Qwen pitch them
Codex is OpenAI's own coding agent, built around OpenAI's models. For anyone still on GPT-5.6 Sol, GPT-5.6 Terra, or GPT-5.4, the Codex docs point to GPT-6 Sol, which they also recommend for complex coding. GPT-6 Sol is also offered in OpenCode, OpenRouter, and GitHub Copilot.
Qwen3.8-Max is offered through OpenRouter and in OpenCode. Qwen says it trained the model with reinforcement learning across harnesses including Claude Code and Codex, and built a self-evolving harness during an autonomous coding run of more than 10 days. It calls Qwen3.8-Max "the most capable model in the Qwen family to date." OpenAI's own line for GPT-6 Sol is that it was "built to power complex coding and agentic workflows."
Both are closed models: Qwen publishes open weights for the base variant Qwen3.8-2.4T-A95B, not for the Max model. EveryToken prices GPT-6 Sol at OpenAI's rates from Codex and OpenCode history, and prices Qwen3.8-Max from OpenRouter's catalog when you use it through OpenRouter, where provider prices can differ from Qwen Cloud's.
Prompt caching
How each maker bills cached tokens
Alibaba Qwen
Implicit caching is automatic and can't be turned off. A hit on Qwen3.8-Max costs $0.25 per million tokens, and written tokens cost ordinary input.
Explicit caching, marked with cache_control, lasts 5 minutes. On Qwen3.8-Max it costs $2.50 per million tokens to create and $0.17 per million to read. This blog prices Qwen with implicit caching.
OpenAI
Prompt caching is on by default. From GPT-5.6 on you can also mark up to four explicit cache breakpoints, while GPT-5.5 and earlier cache automatically only.
From GPT-5.6 on, a cache write costs 1.25x the uncached input price and a cache hit costs 0.1x. GPT-5.5 and earlier add no charge for writing the cache: written tokens are billed as ordinary input, and a hit costs 0.1x on the models compared here.
On GPT-5.6 and later, a cached prefix stays reusable for at least 30 minutes after its last use, and caching starts at 1,024 input tokens.
Source: OpenAI: Prompt caching
Your own numbers
See what Qwen3.8-Max and GPT-6 Sol really cost you.
everyaitoken reads your OpenRouter, Codex, and OpenCode history on your Mac and prices every request at API rates, with what caching saved or cost. $9 once.
FAQ
Questions
Is Qwen3.8-Max cheaper than GPT-6 Sol?
Yes on most work, by a margin that grows with output. Input costs the same $2 per million, output is $6 against $10, and the example agentic session costs $1.80 against $2.10. GPT-6 Sol charges less per automatic cache hit, $0.20 against $0.25.
Is Qwen3.8-Max in Codex's model list?
This page's sources don't cover it. Codex is built around OpenAI's models, though Qwen says it trained Qwen3.8-Max across the Codex harness. Qwen3.8-Max is offered through OpenRouter, in OpenCode, and on Qwen Cloud.
Which model charges more for long prompts?
GPT-6 Sol does. A prompt beyond 272K input tokens costs 2x for input and cache and 1.5x for output across the entire request, and the Qwen3.8-Max rates here list no long-context tier.
Does either model have open weights?
No. GPT-6 Sol is closed, and so is the Qwen3.8-Max API model. Qwen publishes open weights for its base variant, Qwen3.8-2.4T-A95B.
Sources
- Qwen Cloud: Qwen3.8-Max
- Qwen: Qwen3.8
- Alibaba Cloud Model Studio: Model pricing
- OpenRouter: Qwen3.8-Max
- OpenCode docs: Zen
- OpenAI: API pricing
- OpenAI docs: GPT-6 Sol
- OpenAI: Introducing GPT-6 Sol and Luna
- OpenAI: API changelog
- Codex docs: Models
- OpenRouter: GPT-6 Sol
- OpenCode docs: Zen
- GitHub Docs: Supported AI models in Copilot
- Alibaba Cloud Model Studio: Context cache
- OpenAI: Prompt caching