📊 Full opportunity report: Unpacking Qwen3.8-Max's AI Capabilities – A Closer Look At The Numbers on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba has confirmed the specifications and benchmark results of its Qwen3.8-Max model, a 2.4 trillion-parameter AI with strong performance metrics. Open weights will be available next week, marking a significant milestone in open-weight models.

Alibaba has officially confirmed the specifications and benchmark results of its Qwen3.8-Max model, a 2.4 trillion-parameter AI, with open weights scheduled for release next week. This marks a significant step in transparency and openness for large-scale models, as the company revealed detailed performance metrics after two weeks of speculation.

On August 3, Alibaba published the full benchmark table for Qwen3.8-Max, previously only previewed in July, confirming its 2.4 trillion total parameters and approximately 95 billion active parameters per query. The model is built on the Qwen3.5 architecture, employs sparse mixture-of-experts, and supports multimodal inputs including text, images, and videos, with text output.

The benchmark table shows the model outperforming several competitors on key tasks: it scored 86.6 on Terminal-Bench 2.1, surpassing Claude Opus 4.8 and Claude Fable 5, but trailing GPT-5.6 Sol at 88.8. On PaperBench, it achieved the top score of 93.0, and it demonstrated strong performance in multimodal and agentic tasks, such as OSWorld-Verified at 86.1 and Parametric CAD Bench at 91.5. Notably, the model excelled in research reproduction and long-horizon agent tasks, beating its predecessor significantly.

However, the model trails on deep software-engineering benchmarks like SWE-bench Pro, where it scored 67.7 against Fable 5’s 80.0, indicating areas for improvement. The active-parameter count clarifies that about 4% of the total network fires per token, a key detail settling prior uncertainties about its size.

At a glance
reportWhen: announced August 3, 2023; details relea…
The developmentAlibaba announced the official specifications and benchmark results of Qwen3.8-Max, confirming its 2.4 trillion parameters and upcoming open weights release.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Impact of Qwen3.8-Max’s Benchmark Results and Open Release

The confirmed specifications and benchmark scores of Qwen3.8-Max demonstrate Alibaba’s capability to build a highly competitive large language model, especially in multimodal and agentic tasks. The upcoming open weights will enable broader access, potentially influencing the AI ecosystem by providing a new open-weight option for research and deployment. The performance in research reproduction and long-horizon tasks indicates progress in AI reasoning and planning, relevant for both industry and academia.

While the model excels in several areas, its lower scores on software engineering benchmarks highlight ongoing challenges. The decision to release open weights next week marks a strategic move toward transparency and community engagement, with implications for AI development standards and competition.

Agentic Spec-Driven Development: A Practical Method for Using AI to Build Complete Specifications for Software, Products, and Knowledge Work

Agentic Spec-Driven Development: A Practical Method for Using AI to Build Complete Specifications for Software, Products, and Knowledge Work

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background and Development of Alibaba’s Large Language Models

Alibaba’s Qwen series has been under development since at least July 2023, with the company initially previewing the model as a stealth project called kaleb during the World AI Conference in Shanghai. The model's parameters and capabilities have been gradually disclosed, culminating in today’s full benchmark reveal. The release follows a trend among Chinese AI firms to challenge Western dominance by building and sharing large, multimodal models.

Previous models like Kimi K3 and other Chinese LLMs have aimed at balancing performance with open access, but Alibaba’s recent move to publish detailed benchmark results and open weights signifies a shift toward more transparent AI development. The model's architecture is based on the Qwen3.5 platform, with enhancements aimed at improving agentic reasoning and multimodal understanding.

"Next week, we will release the open weights of Qwen3.8-Max, enabling broader research and deployment opportunities."

— Alibaba spokesperson

Amazon

large language model open weights

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About Open Weights and Real-World Performance

While Alibaba confirmed the upcoming release of open weights next week, details about licensing, licensing terms, and the exact deployment options remain unpublished. It is also unclear how the model’s performance will translate outside benchmark settings, especially in real-world applications. Furthermore, the impact of the open weights on the broader AI ecosystem and competition is still to be seen.

Amazon

multimodal AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Alibaba’s Qwen3.8-Max and Community Engagement

Alibaba plans to release the open weights of Qwen3.8-Max next week, likely accompanied by documentation and licensing details. The community will then evaluate its performance in practical deployments and research. Additionally, further benchmark results and potential updates to the model’s capabilities are expected in the coming months, shaping its role in AI development.

Evals for AI Engineers: Systematically Measuring and Improving AI Applications

Evals for AI Engineers: Systematically Measuring and Improving AI Applications

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

When will Alibaba release the open weights for Qwen3.8-Max?

Alibaba has announced that the open weights will be available next week, with exact date to be confirmed.

How does Qwen3.8-Max compare to other large language models?

It outperforms several competitors on key benchmarks like Terminal-Bench 2.1 and PaperBench, but trails on software engineering tasks such as SWE-bench Pro. Its multimodal and agentic capabilities are notably strong.

What are the main limitations of Qwen3.8-Max based on current benchmarks?

The model scores lower on software engineering benchmarks like SWE-bench Pro and FrontierSWE, indicating ongoing challenges in deep coding tasks.

Will the open weights be fully open-source?

The licensing details are still unpublished, but historically Alibaba’s open models have used Apache 2.0 licenses. The upcoming release will clarify the licensing terms.

Source: ThorstenMeyerAI.com

You May Also Like

How ChatGPT Can Streamline Your Blog Writing

Discover how ChatGPT can revolutionize your blog writing process and unlock new levels of productivity—here’s what you need to know.

Three Public Vulnerabilities. Chained.

A chain of three known vulnerabilities was exploited in the TanStack npm packages on May 11, 2026, leading to widespread compromise. Details reveal public research was weaponized rapidly.

The MiniMax H3 AI Transformer Ships With Sound — But What Does ‘Open’ Signify?

MiniMax launched H3 on July 31, 2026, offering 2K video with synchronized sound via a novel joint prediction architecture, but ‘open’ refers only to base model weights, not full open source.

Understanding AI’s Management Challenges Post-Accurate Output

Analysis of AI’s ability to translate correct analysis into trustworthy, completed work under real-world pressures, based on recent experiments.