Alibaba shipped its largest model ever this morning and forgot the scoreboard. Qwen3.8-Max arrived with a parameter count (2.4 trillion), a price ($2.00 per million input tokens, $6.00 output, $0.25 implicit caching), a promise that open weights for both Qwen3.8-Max and Qwen3.8-27B land next week, and a tagline calling it "a new bar for coding and cowork." What it did not arrive with was a single benchmark number. No SWE-bench Verified. No SWE-bench Pro. No Terminal-Bench. No methodology section, no eval config, no table at all.
That absence is the story, because this is now the second Qwen3.8 launch event in three weeks that has skipped the numbers. The July 19 preview drew the same complaint, and Alibaba's answer then was that the finished release would settle it. The finished release is here and it did not.
RelatedAnthropic Ships Claude Opus 5 at Half of Fable 5’s Price
What did Alibaba actually claim?
The announcement is built on descriptions rather than measurements. Qwen3.8-Max is pitched at long-horizon autonomous work: Alibaba says it ran "10+ days of self-evolving development, from empty folder to production without hand-holding," and points to a complete project trace on GitHub as evidence. It describes system-level autonomous planning, closed-loop adaptive learning, and native multimodal intelligence in which vision acts as a continuous feedback signal for planning and self-correction. The company also says the model produces production-quality deliverables across hundreds of real office and analysis tasks, which is where the "cowork" half of the tagline comes from.
A published project trace is a real artifact and it is worth more than nothing. It is also not a benchmark. A trace shows one run, chosen by the vendor, on a task chosen by the vendor, with no baseline to compare against and no way for anyone else to reproduce the conditions. It answers "can this model do a thing once" and not "how often does this model do the thing, versus the alternatives, under identical conditions." Those are different questions and only the second one helps you pick a model.
Why does a missing score matter this much?
Because Alibaba's last flagship gave us a measurement of how far its self-reported numbers drift. Qwen3.7 Max shipped in May 2026 with a claimed 80.4% on SWE-bench Verified. When vals.ai ran the same model itself on a neutral bash-only harness, it came out at 68.8%. That 11.6-point gap is the widest vendor-versus-independent spread on our AI Coding Leaderboard, and it is not an accusation of dishonesty: SWE-bench scores the model together with its scaffolding, so a vendor running its own tuned harness will legitimately post a higher number than a neutral one. The gap is a units problem, not a fraud problem. It still means an unverified vendor figure is not interchangeable with a measured one.
The more uncomfortable detail sits one layer down. On vals.ai's harness, Alibaba's newest Max-class model scores lower than its own previous generation. Qwen 3.7 Max measures 68.8%, while Qwen 3.6 Plus lands at 73.4%, Qwen 3.6 Max Preview at 72.8%, and the small open-weight Qwen 3.6 27B at 70.0%. Read against the vendor's 80.4% claim, the newest model looks like a clear step up. Read against a neutral harness, it is the weakest recent Alibaba entry on the board.
How does it compare on what we can actually check?
Price, context and licensing are all verifiable today, so those are worth lining up even while the capability question stays open. Against the models it is implicitly aiming at, Qwen3.8-Max is priced in the middle of the pack and is the only one in this group with no confirmed score of any kind.
| Qwen3.8-Max | Qwen3.7 Max | Kimi K3 | DeepSeek V4 Flash | |
|---|---|---|---|---|
| SWE-bench Verified | none published | 80.4% vendor / 68.8% measured | 93.4% measured | none published |
| Price in / out per 1M | $2.00 / $6.00 | $1.25 / $3.75 | varies by host | $0.14 / $0.28 |
| Context | 1M | 1M | 256K | 1M |
| Open weights | promised next week | no | yes | yes (MIT) |
| Rankable today | no | yes | yes | no |
What it means for the market
Alibaba (NYSE: BABA, HK: 9988) has spent the past year arguing that its model line is close enough to the US frontier to justify the capex behind it, and Qwen has been the centrepiece of that argument. Launching the biggest model in the family without a scoreboard weakens the pitch at exactly the moment it needs to be strongest, because it leaves the claim resting on a narrative rather than a comparable. The signal for investors is not that the model is bad, nobody knows yet, but that the disclosure discipline is inconsistent, and inconsistent disclosure is the thing that gets discounted when a Chinese lab's numbers are compared against a US lab's.
RelatedLaguna S 2.1: 8B Active Params, 70% Terminal-Bench
The open-weights promise is the part with real strategic weight. If a Max-class 2.4T model genuinely ships under an open licence next week, that is a first for the tier and it pressures the economics of every closed API priced above $2 per million input tokens. Watch whether the weights actually appear on schedule and under what licence, because a delayed or restricted release would be the more informative event than the launch itself.
Our take
We put Qwen3.8-Max on the AI Coding Leaderboard today as an unranked row, alongside five other models sitting in the same queue. That is not a snub. It is the only honest position available when a vendor publishes no number, because the alternative is inventing a placement out of adjectives. A model with a public price and a private capability profile is a purchase you cannot evaluate, and "a new bar for coding" is a sentence, not a measurement.
The fix here is cheap and entirely in Alibaba's hands. Publish the SWE-bench Verified score, name the harness, and let vals.ai or anyone else run it independently. Qwen has a genuinely strong open-weight track record and the 27B models are widely liked by people who self-host. That reputation is what makes the silence odd rather than expected.
- Do the weights ship on time? Alibaba said next week for both Qwen3.8-Max and Qwen3.8-27B. A slip, or a licence with commercial restrictions, would say more about the release than the launch post did.
- Does a benchmark table appear after the fact? A score published a week late still counts, and would move the model from unranked to ranked on our board immediately.
- Where does an independent harness land it? Given the 11.6-point spread on Qwen3.7 Max, expect a neutral result well below whatever Alibaba eventually claims.
- Does the GitHub trace hold up? The 10-day autonomous build is checkable by anyone willing to read it. If it is as clean as described, that is a genuinely novel artifact even without a benchmark.
- OfficialQwen3.8-Max: A New Bar for Coding and Cowork , the launch announcement, August 3, 2026
- Benchmarkvals.ai SWE-bench Verified , independent bash-only harness, 75 systems, updated July 31, 2026
- ReferenceGenZTech AI Coding Leaderboard , where Qwen3.8-Max now sits unranked pending a score
- ReferenceAI coding cost calculator , model the $2.00/$6.00 pricing against your own token volume
Original analysis by GenZTech, built on Alibaba's launch materials and vals.ai's independent evaluations.
