/00 — boot sequence

Hello.

Article

Qwen3.8-Max: Open Weights Coming, Top Agentic Index

August 9, 2026•7 min read
Qwen AI Models Open Source AI Agents Benchmarks Alibaba

Alibaba officially released Qwen3.8-Max on August 3, 2026, and the launch has shaken up the frontier AI field. The 2.4 trillion parameter mixture-of-experts model is the most capable in the Qwen family so far, it now tops Artificial Analysis' agentic index, and for the first time Alibaba says the weights of a Max-class model will be open-sourced. At $2 per million input tokens it also undercuts most of the competition while posting headline benchmark wins such as 86.1 on OSWorld-Verified and 93.0 on PaperBench.

What Qwen3.8-Max is

Qwen3.8-Max is a multimodal MoE model with 2.4 trillion total parameters and about 95 billion active per token, built on the architecture of Qwen 3.5. Alibaba positions it as an autonomous coworker rather than a chat model: it is designed to run multi-day coding projects, reproduce research papers, operate computer interfaces, and iterate on its own output over hundreds of turns.

It is available today through QwenCloud with an OpenAI-compatible API and an Anthropic-compatible endpoint, so it plugs into Claude Code and most agent harnesses without new tooling. The API exposes a reasoning_effort parameter (xhigh, medium, low) to trade depth against cost, and preserve_thinking is on by default.

The detail developers care about most: the release post says open weights will land next week, together with Qwen3.8-27B. This would be the first time a Qwen-Max-class model can be self-hosted. Alibaba has not disclosed the license, which matters because Moonshot's Kimi K3 showed that "open" can still mean a custom license with a commercial clause for model-as-a-service.

Where it wins on benchmarks

The company published a wide benchmark table against Claude Opus 4.8, Claude Fable 5, GPT-5.6 Sol (max), and its own Qwen3.7-Max. Selected numbers:

BenchmarkQwen3.8-MaxGPT-5.6 Sol MaxClaude Fable 5Claude Opus 4.8
OSWorld-Verified86.183.285.083.4
PaperBench93.090.588.880.3
Terminal Bench 2.186.688.884.684.6
SWE-bench Pro67.764.680.069.2
Agents' Last Exam (score)52.453.6n/a45.1

Qwen3.8-Max leads on PaperBench, OSWorld-Verified, and LVBench (81.8), while it trails Fable 5 on SWE-bench Pro and GPT-5.6 Sol on Terminal Bench 2.1 and Agents' Last Exam pass rate. In other words: best in class at some agentic tasks, competitive but not dominant elsewhere.

Independent trackers mostly agree. Artificial Analysis put the model on its agentic index as the best overall model on August 6, and scored it 56 on the Intelligence Index, level with Claude Opus 4.8 and a point behind Kimi K3 (57). On GDPval-AA, its work-oriented benchmark, Qwen jumps to 1,739 Elo, passing Kimi K3 (1,685); only Claude Opus 5 (1,852) scores higher.

The autonomous coding demonstrations

Alibaba's demonstration runs are worth reading even with a skeptical hat on. In one, Qwen3.8-Max built the oh-my-cli project over a 10+ day autonomous run: as of July 30 the repository had 265 commits, 127 pull requests, and 151 issues, with the model claiming, dispatching, implementing, and self-checking its own backlog through a GitHub-issue state machine.

In another, the model reproduced a research paper on data selection for LLM reasoning from scratch: no starter code, just the paper and GPUs. It wrote roughly 7,600 lines of code, ran 33 rounds of training, and after about 125 hours beat the paper's own method by 2.7 points on AIME24. It also entered a Tianchi competition with 526 human teams and finished ahead of 458 of them (87% of the field), climbing from 0.60 to 0.853 accuracy across 45 submissions.

These are vendor-run demos, not independent reproductions, and VentureBeat flags the same caveat. Treat them as upper bounds. The pattern they describe, a model that keeps improving its own approach over hundreds of turns, matches what the benchmark tables show.

Pricing and the cost per task

Qwen3.8-Max launches at $2 per million input tokens and $6 per million output, with cache hits at $0.25. That is a 20% cut on input versus Qwen3.7-Max ($2.50/$7.50), and it lands far below the US flagships:

ModelInput ($/1M)Output ($/1M)
Qwen3.8-Max2.006.00
GLM-5.21.404.40
Kimi K33.0015.00
GPT-5.6 Sol (standard)5.0030.00
Claude Opus 55.0025.00
Claude Fable 510.0050.00

Combined in/out pricing works out to less than a quarter of GPT-5.6 Sol Max and less than a third of Claude Opus 5.

The caveat is efficiency. Artificial Analysis' cost-per-task model puts Qwen3.8-Max at $1.14 per Intelligence Index task, more than double Qwen3.7 Max ($0.53) and above Kimi K3 ($0.86) and GLM-5.2 ($0.57). The model takes 64 steps per task where Kimi K3 takes 14, and it resends conversation history each step, inflating input tokens 15x. It also regressed on two AA measures: AA-Omniscience fell 10 points and the hallucination rate went from 23% to 40%. Thorough work has a price, and part of that price is guessing when it should say "I don't know".

Open weights and the business model question

Two open questions dominate the next week. First, the license: Alibaba has not said whether Qwen3.8-Max will ship under Apache 2.0 or a custom license. That single decision decides whether enterprises can legally fine-tune and self-host without commercial terms. Second, the business model: Reuters reported on August 7 that Alibaba is exploring charging its biggest users of the next open-source Qwen model, with revenue sharing for large deployments. Alibaba has not confirmed the plan, but the direction is clear. "Open weights" no longer guarantees "free for everyone at any scale", and teams building products on Qwen should watch the license terms when the weights land.

What this means for developers

You can try the model today. The API is OpenAI-compatible, so the change is a base URL and a key:

python

For Claude Code, point ANTHROPIC_BASE_URL at the Anthropic-compatible endpoint and set ANTHROPIC_MODEL=qwen3.8-max.

Three practical reads:

  1. If you build agent pipelines, benchmark on your own workloads before switching. The model is strong at long-horizon autonomy but slower and more token-hungry per task than Kimi K3.
  2. If you self-host, wait for the license before committing. The weights drop this week, but "open weights" and "open license" are not the same thing.
  3. If you are price-sensitive, GLM-5.2 and Kimi K3 remain cheaper per task; Qwen3.8-Max wins on raw frontier performance at $2/$6.

Frequently Asked Questions

When do the Qwen3.8-Max weights release? Alibaba's August 3 announcement says next week, together with Qwen3.8-27B.

What license will the open weights use? Not disclosed as of August 9. Check the Hugging Face and ModelScope releases before planning commercial use.

What does Qwen3.8-Max cost? $2 per million input tokens, $6 per million output, and $0.25 for cache hits via QwenCloud.

How does it compare with Kimi K3? One point behind on the AA Intelligence Index (56 vs 57) but ahead on GDPval-AA (1,739 vs 1,685). Kimi K3 is cheaper per task ($0.86 vs $1.14).

Can I run it locally? Once the weights land, yes in principle: it is a 2.4T MoE with 95B active parameters, so a single high-end GPU can serve it with quantization, though that is a serious deployment either way.

Key Takeaways

  • Qwen3.8-Max launched August 3 as Alibaba's most capable model: 2.4T total, 95B active, multimodal.
  • It leads OSWorld-Verified (86.1), PaperBench (93.0), and the AA agentic index; it trails on SWE-bench Pro and cost per task.
  • Pricing is $2/$6 per million tokens, undercutting GPT-5.6 Sol Max by more than 4x combined.
  • Open weights arrive next week with Qwen3.8-27B, but the license is still unknown.
  • Reuters reports Alibaba may charge heavy users of the next open-source Qwen model, a shift in how "open" models make money.

Conclusion

Qwen3.8-Max is the strongest open-weight signal of the summer. It matches or beats closed frontier models on several agentic benchmarks, it is aggressively priced, and the open-weights release, if the license cooperates, gives developers a self-hostable flagship that did not exist before. The license and the cost-per-task numbers will decide how much of that promise is real. Watch the weight release this week.


Sources: Alibaba Qwen blog: Qwen3.8-Max launch post, VentureBeat: Qwen3.8-Max arrives with a bold claim, The Decoder: Qwen3.8 Max catches Claude Opus 4.8 but Kimi K3 still scores higher, Artificial Analysis: agentic index, Reuters via The Next Web: Alibaba wants to charge the biggest users of its open AI model, Hacker News discussion

Automated Transmission

This entry was synthesized and populated dynamically using native API integrations.

Resources & Links