📊 Full opportunity report: Unpacking Qwen3.8-Max's AI Capabilities – A Closer Look At The Numbers on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Alibaba has confirmed the specifications and benchmark results of its Qwen3.8-Max model, a 2.4 trillion-parameter AI with strong performance metrics. Open weights will be available next week, marking a significant milestone in open-weight models.
Alibaba has officially confirmed the specifications and benchmark results of its Qwen3.8-Max model, a 2.4 trillion-parameter AI, with open weights scheduled for release next week. This marks a significant step in transparency and openness for large-scale models, as the company revealed detailed performance metrics after two weeks of speculation.
On August 3, Alibaba published the full benchmark table for Qwen3.8-Max, previously only previewed in July, confirming its 2.4 trillion total parameters and approximately 95 billion active parameters per query. The model is built on the Qwen3.5 architecture, employs sparse mixture-of-experts, and supports multimodal inputs including text, images, and videos, with text output.
The benchmark table shows the model outperforming several competitors on key tasks: it scored 86.6 on Terminal-Bench 2.1, surpassing Claude Opus 4.8 and Claude Fable 5, but trailing GPT-5.6 Sol at 88.8. On PaperBench, it achieved the top score of 93.0, and it demonstrated strong performance in multimodal and agentic tasks, such as OSWorld-Verified at 86.1 and Parametric CAD Bench at 91.5. Notably, the model excelled in research reproduction and long-horizon agent tasks, beating its predecessor significantly.
However, the model trails on deep software-engineering benchmarks like SWE-bench Pro, where it scored 67.7 against Fable 5’s 80.0, indicating areas for improvement. The active-parameter count clarifies that about 4% of the total network fires per token, a key detail settling prior uncertainties about its size.
For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.
▲ All performance figures: Alibaba’s own harnessThe claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.
“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.
“Qwen3.8 is going open-weight” describes three things with very different deployment realities.
OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.
A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.
The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.
Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.
- The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
- More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
- If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
- The 27B sibling could become the best local agent model on hardware people already own.
- Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
- The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
- “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
- Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
and it says “second only” depends entirely on which row you read.
Impact of Qwen3.8-Max’s Benchmark Results and Open Release
The confirmed specifications and benchmark scores of Qwen3.8-Max demonstrate Alibaba’s capability to build a highly competitive large language model, especially in multimodal and agentic tasks. The upcoming open weights will enable broader access, potentially influencing the AI ecosystem by providing a new open-weight option for research and deployment. The performance in research reproduction and long-horizon tasks indicates progress in AI reasoning and planning, relevant for both industry and academia.
While the model excels in several areas, its lower scores on software engineering benchmarks highlight ongoing challenges. The decision to release open weights next week marks a strategic move toward transparency and community engagement, with implications for AI development standards and competition.

Agentic Spec-Driven Development: A Practical Method for Using AI to Build Complete Specifications for Software, Products, and Knowledge Work
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background and Development of Alibaba’s Large Language Models
Alibaba’s Qwen series has been under development since at least July 2023, with the company initially previewing the model as a stealth project called kaleb during the World AI Conference in Shanghai. The model's parameters and capabilities have been gradually disclosed, culminating in today’s full benchmark reveal. The release follows a trend among Chinese AI firms to challenge Western dominance by building and sharing large, multimodal models.
Previous models like Kimi K3 and other Chinese LLMs have aimed at balancing performance with open access, but Alibaba’s recent move to publish detailed benchmark results and open weights signifies a shift toward more transparent AI development. The model's architecture is based on the Qwen3.5 platform, with enhancements aimed at improving agentic reasoning and multimodal understanding.
"Next week, we will release the open weights of Qwen3.8-Max, enabling broader research and deployment opportunities."
— Alibaba spokesperson
As an affiliate, we earn on qualifying purchases.
Remaining Questions About Open Weights and Real-World Performance
While Alibaba confirmed the upcoming release of open weights next week, details about licensing, licensing terms, and the exact deployment options remain unpublished. It is also unclear how the model’s performance will translate outside benchmark settings, especially in real-world applications. Furthermore, the impact of the open weights on the broader AI ecosystem and competition is still to be seen.
As an affiliate, we earn on qualifying purchases.
Next Steps for Alibaba’s Qwen3.8-Max and Community Engagement
Alibaba plans to release the open weights of Qwen3.8-Max next week, likely accompanied by documentation and licensing details. The community will then evaluate its performance in practical deployments and research. Additionally, further benchmark results and potential updates to the model’s capabilities are expected in the coming months, shaping its role in AI development.

Evals for AI Engineers: Systematically Measuring and Improving AI Applications
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
When will Alibaba release the open weights for Qwen3.8-Max?
Alibaba has announced that the open weights will be available next week, with exact date to be confirmed.
How does Qwen3.8-Max compare to other large language models?
It outperforms several competitors on key benchmarks like Terminal-Bench 2.1 and PaperBench, but trails on software engineering tasks such as SWE-bench Pro. Its multimodal and agentic capabilities are notably strong.
What are the main limitations of Qwen3.8-Max based on current benchmarks?
The model scores lower on software engineering benchmarks like SWE-bench Pro and FrontierSWE, indicating ongoing challenges in deep coding tasks.
Will the open weights be fully open-source?
The licensing details are still unpublished, but historically Alibaba’s open models have used Apache 2.0 licenses. The upcoming release will clarify the licensing terms.
Source: ThorstenMeyerAI.com