AI-generated content. This article was researched and written by an automated AI editorial system and published without prior human review. Every factual claim is checked against cited primary sources before publication, but no journalist read this page before you did — treat it accordingly, and report anything that looks wrong. How this works ›

Some links on this page are affiliate links. We may earn a commission at no extra cost to you.
Updated: Aug 4, 2026
·
alibabaqwenopen-weightsbenchmarkschinapricing

Alibaba says Qwen 3.8 Max beats GPT-5.6 Sol. Independent evals put it 10th — and you may not be licensed to use it.

TL;DR: Alibaba shipped Qwen 3.8 Max on 3 August 20262.4 trillion total parameters with roughly 95B active per token, a 1M-token context, 128k max output, multimodal, at $2 per million input / $6 per million output ($0.25/M cached). A smaller Qwen 3.8-27B was announced as open-weight, with Max weights promised “next week” on Hugging Face and ModelScope. Alibaba’s own benchmarks put it ahead of GPT-5.6 Sol and Claude Opus 4.8 on PaperBench (93.0) and IFBench (82.8 vs 72.7). Independent numbers arrived within a day and read differently: 10th overall and 2nd among open models on the Vals Index (66.1), 4th on Frontend Code Arena, 2nd on Vision Arena, SWE-bench 87.3%. Both readings are true. Two things deserve more attention than the leaderboard: evaluators noted Alibaba modified benchmark timeouts versus standard settings, and the open-weight licence reportedly carries prohibitions covering the US, EU, UK and Korea — unclarified by Alibaba.

What shipped

Qwen 3.8 Max is Alibaba’s new flagship, and on paper it is a serious piece of engineering.

2.4 trillion total parameters, with approximately 95 billion active per token — an activation ratio around 4%, which is what makes a model this size servable at consumer-grade API prices. A 1,000,000-token context window with 128k maximum output. Multimodal input, with native visual feedback inside planning and coding loops rather than bolted on as a separate vision call.

Pricing on QwenCloud: $2 per million input tokens, $6 per million output, $0.25 per million cached. For context, that is roughly a fifth of what Claude Opus 5 costs on output, and it undercuts most Western frontier pricing by a wide margin — though it sits well above DeepSeek V4-Flash at $0.14/$0.28.

Alibaba also announced Qwen 3.8-27B as an open-weight release, with Max weights promised roughly a week later on Hugging Face and ModelScope.

Two true stories about the same model

Alibaba’s version. PaperBench at 93.0, outpacing both GPT-5.6 Sol and Claude Opus 4.8. IFBench at 82.8 against Sol’s 72.7. Framed as near-frontier performance at a fraction of frontier pricing.

The independent version, which arrived within about a day:

EvaluationQwen 3.8 MaxNotes
Vals Index (overall)66.1 — 10th overall, 2nd among open modelsIndependent aggregate
Frontend Code Arena#4 (1,668 Elo)Human preference
Vision Arena#2 (1,305 Elo)Human preference
SWE-bench87.3%Ahead of GPT-5.5 at 82.6%
Terminal-Bench 2.167.4

Neither picture is dishonest. They are measurements taken under different conditions, and the headline framings diverge because vendors choose the benchmarks where they lead — which is not a Chinese-lab phenomenon, it is a launch-day phenomenon.

The useful reading is the second one, and it is genuinely good news. Tenth overall and second among open models, at $2/$6, is a strong result. So is the trajectory: Qwen 3.7 Max scored 57.5 on the Vals Index; 3.8 scores 66.1. An 8.6-point gain in about two and a half months says more about where this is heading than any single leaderboard position.

The timeout detail, and why it matters more than it sounds

Vals noted something specific: Alibaba’s reported results modified benchmark timeouts relative to standard settings.

That is not fabrication, and it is worth being careful about the accusation. But on agentic benchmarks — the ones measuring whether a model can complete multi-step tasks — a longer timeout lets the model try more approaches, backtrack further, and recover from more mistakes. Scores rise without any change in underlying capability. It is one of the easiest ways for two honest parties to produce different numbers for the same model.

This is now the fourth consecutive major release where vendor benchmark claims required an asterisk:

There is an encouraging difference here, though. In this case independent evaluation arrived within a day, from multiple sources, with methodology notes attached. That is the ecosystem working. It is also the strongest available argument for the point made in Monday’s coverage of Astra’s formally verified proofs: a claim someone else can check is worth categorically more than a claim you must take on trust.

The licence is the thing to actually worry about

Buried under the benchmark discussion is the detail most likely to matter to anyone reading this.

Evaluators examining the licence terms flagged apparent prohibitions covering the United States, the European Union, the United Kingdom and Korea. Alibaba has not publicly clarified. The same concern was raised recently about MiniMax H3, so it may reflect an emerging pattern in Chinese open-weight licensing rather than a one-off.

If those terms hold as written, the consequence is blunt: for most readers of this article, the weights would be legally unusable for commercial deployment regardless of how good the model is. A licence prohibiting the US, EU and UK excludes the large majority of the Western commercial market.

That reframes the release. “Open weights” is doing real work in the marketing, and it may turn out to describe a model you can look at but not deploy. When the files land, read the licence before the benchmark table. That ordering is not usually the right advice; here it is.

It also sits inside a policy environment moving the same direction from the other side. US proposals on open-source model curbs and chip provisions would restrict Western use of Chinese open weights by statute. If the licence restricts it by contract simultaneously, the practical availability of these models in Western production stacks narrows from both ends at once.

The hardware wall is unchanged

Loading Qwen 3.8 Max reportedly requires more than 1TB of memory and at least eight H100 or B300-class GPUs.

This is the same story as Kimi K3: the licence may say you are free, the memory requirement says you are renting. Nobody is running a 2.4-trillion-parameter model on a workstation.

The Qwen 3.8-27B is the one worth attention for actual local deployment, and it is telling that it received a fraction of the coverage. For most teams the realistic options are the API, a hosted provider, or the small model — not self-hosting the flagship.

Why this matters

The value tier keeps compressing, and it is now crowded. DeepSeek V4-Flash, Kimi K3, GLM-5.2 and now Qwen 3.8 Max occupy a band that did not exist eighteen months ago: agent-capable, open-weight or near it, priced where the bill stops driving architecture decisions. Western labs have no equivalent at this price point.

Trajectory beats position. 57.5 to 66.1 in ten weeks is the number to hold onto. Ranking tenth today matters less than closing eight points in a quarter, and Alibaba has now done this repeatedly since the 3.6 Max preview in April.

The verification ecosystem is maturing faster than the marketing. A year ago vendor benchmarks stood unchallenged for weeks. This one was independently contextualised inside 24 hours, with the methodology difference named. That is a real improvement in the information environment, and it should change how much weight a buyer gives to launch-day claims — which is now close to none, because better numbers arrive almost immediately.

Licence risk is becoming the real differentiator in open weights. Capability among the leading Chinese open-weight models is converging. What separates them for a Western buyer increasingly is not quality but whether you are permitted to use them — a question benchmark tables cannot answer.

Honest caveats

The Max weights do not exist yet. “Next week” is an intention. Until files are on Hugging Face, this is a proprietary API model with an open-weight promise attached.

The licence reading is preliminary. It comes from evaluators examining the terms, not from an Alibaba statement or a legal opinion, and Alibaba has not clarified. It could be narrower than it appears, or amended before release. Treat it as a serious flag to verify, not a settled fact.

Independent coverage is early. The Vals Index and Arena placements are the first substantial third-party numbers and will move as more evaluations land. Arena Elo in particular is unstable in the days after a release.

The vendor and independent numbers are not directly comparable because of the timeout difference, so the gap between “beats Sol” and “10th overall” is partly a methodology artefact rather than purely a marketing one.

Data residency is unchanged. API traffic goes to Alibaba infrastructure. For work you would not send to a Chinese jurisdiction, the API is not the answer and the weights may not be legally available either.

What to do

If you are evaluating a value-tier model this week: Qwen 3.8 Max belongs on the shortlist, on independent numbers rather than Alibaba’s. Run your own tasks — that advice is now mandatory rather than cautious, given how much of the gap here comes from evaluation conditions.

If you were waiting for the open weights: wait a little longer, and read the licence first. If the geographic prohibitions hold, the release is not usable for Western commercial deployment and the benchmark discussion is moot for you.

If you want local deployment: look at Qwen 3.8-27B, not Max. The flagship needs a rack.

If you are on ChatGPT or Claude for frontier work: nothing here displaces them. Qwen sits tenth overall on independent aggregate. What it changes is the economics of the high-volume, lower-stakes half of a workload — the same argument that applies to DeepSeek, now with a second credible option.


Related: DeepSeek’s cheap tier beat its own flagship · Kimi K3’s independent benchmarks · Astra’s ten proofs and the verification lesson

Frequently asked questions

Is Qwen 3.8 Max actually better than GPT-5.6 Sol?

On Alibaba's own evaluations it leads on several benchmarks, including PaperBench and IFBench. On independent evaluations it does not: the Vals Index places it 10th overall and 2nd among open models. Both statements are accurate — they measure different things under different conditions. The independent placement is the one to plan around, and 10th overall for a model at $2/$6 per million tokens is a genuinely strong result.

Can I download the weights?

Not yet for the Max model. Alibaba said open weights would follow roughly a week after the 3 August announcement, on Hugging Face and ModelScope. The smaller Qwen 3.8-27B was announced as open-weight. Until files exist, treat this as an announced intention rather than a shipped release.

What is the licence problem?

Evaluators examining the licence terms flagged apparent prohibitions covering the United States, the EU, the UK and Korea. Alibaba has not publicly clarified the point. If those restrictions hold as written, the weights would be unusable for commercial deployment by most readers of this article regardless of quality — which makes the licence, not the benchmark table, the first thing to check when files appear.

Can I run it locally?

No. Loading the Max model reportedly requires more than 1TB of memory and at least eight H100 or B300-class GPUs. As with Kimi K3, the licence may say you are free while the hardware bill says you are renting. The 27B model is the one with realistic local deployment potential.

What did the timeout issue involve?

The evaluation group Vals noted that Alibaba's reported results modified benchmark timeouts relative to standard settings. Longer timeouts let a model spend more time per task, which can raise scores on agentic benchmarks without any underlying capability change. It does not invalidate the results, but it does mean vendor and independent numbers are not directly comparable.

Sources

Related tool reviews

Questions or corrections? Email Pick Right. Want the full list? See all news.