Alibaba shipped its largest model ever this morning and forgot the scoreboard. Qwen3.8-Max arrived with a parameter count (2.4 trillion), a price ($2.00 per million input tokens, $6.00 output, $0.25 implicit caching), a promise that open weights for both Qwen3.8-Max and Qwen3.8-27B land next week, and a tagline calling it "a new bar for coding and cowork." What it did not arrive with was a single benchmark number. No SWE-bench Verified. No SWE-bench Pro. No Terminal-Bench. No methodology section, no eval config, no table at all.

That absence is the story, because this is now the second Qwen3.8 launch event in three weeks that has skipped the numbers. The July 19 preview drew the same complaint, and Alibaba's answer then was that the finished release would settle it. The finished release is here and it did not.

RelatedAnthropic Ships Claude Opus 5 at Half of Fable 5’s Price

Qwen3.8-Max launch card: published specs versus missing evaluation dataA two-column comparison. The left column lists five specifications Alibaba published: 2.4 trillion parameters, $2.00 and $6.00 per million tokens, 1 million token context, open weights promised, and preview API access. The right column lists five evaluation fields left blank: SWE-bench Verified, SWE-bench Pro, Terminal-Bench, harness methodology, and any independent evaluation.QWEN3.8-MAX LAUNCH CARDPublishedParameters2.4TPrice in / out$2.00 / $6.00Context1MOpen weightsnext weekAPI accesspreviewLeft blankSWE-bench Verified--SWE-bench Pro--Terminal-Bench--Harness / method--Independent eval--A model you can price to the cent and cannot rank at all.genztech.blog
Fig 1 Everything Alibaba quantified at launch, and everything it did not. Compiled from the Qwen3.8-Max announcement, August 3, 2026.

What did Alibaba actually claim?

The announcement is built on descriptions rather than measurements. Qwen3.8-Max is pitched at long-horizon autonomous work: Alibaba says it ran "10+ days of self-evolving development, from empty folder to production without hand-holding," and points to a complete project trace on GitHub as evidence. It describes system-level autonomous planning, closed-loop adaptive learning, and native multimodal intelligence in which vision acts as a continuous feedback signal for planning and self-correction. The company also says the model produces production-quality deliverables across hundreds of real office and analysis tasks, which is where the "cowork" half of the tagline comes from.

A published project trace is a real artifact and it is worth more than nothing. It is also not a benchmark. A trace shows one run, chosen by the vendor, on a task chosen by the vendor, with no baseline to compare against and no way for anyone else to reproduce the conditions. It answers "can this model do a thing once" and not "how often does this model do the thing, versus the alternatives, under identical conditions." Those are different questions and only the second one helps you pick a model.

Why does a missing score matter this much?

Because Alibaba's last flagship gave us a measurement of how far its self-reported numbers drift. Qwen3.7 Max shipped in May 2026 with a claimed 80.4% on SWE-bench Verified. When vals.ai ran the same model itself on a neutral bash-only harness, it came out at 68.8%. That 11.6-point gap is the widest vendor-versus-independent spread on our AI Coding Leaderboard, and it is not an accusation of dishonesty: SWE-bench scores the model together with its scaffolding, so a vendor running its own tuned harness will legitimately post a higher number than a neutral one. The gap is a units problem, not a fraud problem. It still means an unverified vendor figure is not interchangeable with a measured one.

The more uncomfortable detail sits one layer down. On vals.ai's harness, Alibaba's newest Max-class model scores lower than its own previous generation. Qwen 3.7 Max measures 68.8%, while Qwen 3.6 Plus lands at 73.4%, Qwen 3.6 Max Preview at 72.8%, and the small open-weight Qwen 3.6 27B at 70.0%. Read against the vendor's 80.4% claim, the newest model looks like a clear step up. Read against a neutral harness, it is the weakest recent Alibaba entry on the board.

Alibaba models on SWE-bench Verified, vendor claim versus independent measurementHorizontal bar chart. Qwen 3.7 Max is claimed by Alibaba at 80.4 percent but measured by vals.ai at 68.8 percent, a gap of 11.6 points. Qwen 3.6 Plus measures 73.4 percent, Qwen 3.6 Max Preview 72.8 percent and Qwen 3.6 27B 70.0 percent, all above the measured score of the newer 3.7 Max. Qwen3.8-Max has no score of either kind.SWE-BENCH VERIFIED · MINI-SWE-AGENT BASH-ONLYQwen 3.7 Maxclaimed80.4%Qwen 3.7 Maxmeasured68.8%Qwen 3.6 Plus73.4%Qwen 3.6 Max Prev72.8%Qwen 3.6 27B70.0%Qwen3.8-Maxno score published, vendor or independentOn a neutral harness, Alibaba's newest Max scores below its own prior generation.genztech.blog
Fig 2 · benchmark Independent scores from vals.ai's mini-swe-agent bash-only harness, updated July 31, 2026. The 80.4% figure is Alibaba's own, run on its own scaffold.

How does it compare on what we can actually check?

Price, context and licensing are all verifiable today, so those are worth lining up even while the capability question stays open. Against the models it is implicitly aiming at, Qwen3.8-Max is priced in the middle of the pack and is the only one in this group with no confirmed score of any kind.

 Qwen3.8-MaxQwen3.7 MaxKimi K3DeepSeek V4 Flash
SWE-bench Verifiednone published80.4% vendor / 68.8% measured93.4% measurednone published
Price in / out per 1M$2.00 / $6.00$1.25 / $3.75varies by host$0.14 / $0.28
Context1M1M256K1M
Open weightspromised next weeknoyesyes (MIT)
Rankable todaynoyesyesno

What it means for the market

Alibaba (NYSE: BABA, HK: 9988) has spent the past year arguing that its model line is close enough to the US frontier to justify the capex behind it, and Qwen has been the centrepiece of that argument. Launching the biggest model in the family without a scoreboard weakens the pitch at exactly the moment it needs to be strongest, because it leaves the claim resting on a narrative rather than a comparable. The signal for investors is not that the model is bad, nobody knows yet, but that the disclosure discipline is inconsistent, and inconsistent disclosure is the thing that gets discounted when a Chinese lab's numbers are compared against a US lab's.

RelatedLaguna S 2.1: 8B Active Params, 70% Terminal-Bench

The open-weights promise is the part with real strategic weight. If a Max-class 2.4T model genuinely ships under an open licence next week, that is a first for the tier and it pressures the economics of every closed API priced above $2 per million input tokens. Watch whether the weights actually appear on schedule and under what licence, because a delayed or restricted release would be the more informative event than the launch itself.

Our take

We put Qwen3.8-Max on the AI Coding Leaderboard today as an unranked row, alongside five other models sitting in the same queue. That is not a snub. It is the only honest position available when a vendor publishes no number, because the alternative is inventing a placement out of adjectives. A model with a public price and a private capability profile is a purchase you cannot evaluate, and "a new bar for coding" is a sentence, not a measurement.

The fix here is cheap and entirely in Alibaba's hands. Publish the SWE-bench Verified score, name the harness, and let vals.ai or anyone else run it independently. Qwen has a genuinely strong open-weight track record and the 27B models are widely liked by people who self-host. That reputation is what makes the silence odd rather than expected.

What to watch · next 30 days
  • Do the weights ship on time? Alibaba said next week for both Qwen3.8-Max and Qwen3.8-27B. A slip, or a licence with commercial restrictions, would say more about the release than the launch post did.
  • Does a benchmark table appear after the fact? A score published a week late still counts, and would move the model from unranked to ranked on our board immediately.
  • Where does an independent harness land it? Given the 11.6-point spread on Qwen3.7 Max, expect a neutral result well below whatever Alibaba eventually claims.
  • Does the GitHub trace hold up? The 10-day autonomous build is checkable by anyone willing to read it. If it is as clean as described, that is a genuinely novel artifact even without a benchmark.
Primary sources

Original analysis by GenZTech, built on Alibaba's launch materials and vals.ai's independent evaluations.