Alibaba launched Qwen3.8-Max on August 3, 2026, a 2.4-trillion-parameter flagship it calls a new bar for coding, and published no SWE-bench, Terminal-Bench, or SWE-bench Pro score to support it. The pricing is concrete at $2 in and $6 out per million tokens; the capability claim is not.
Read the full story: Qwen3.8-Max lands at 2.4T with no benchmark table →
Transcript
Alibaba shipped its largest model ever this morning and forgot the scoreboard. Qwen3.8-Max arrives at two point four trillion parameters, two dollars per million tokens in, six dollars out, and a promise that open weights land next week. What it does not arrive with is a single benchmark number. No SWE-bench Verified. No SWE-bench Pro. No Terminal-Bench. No methodology at all. The headline claim is that it ran ten days of self-evolving development, from an empty folder to production, with nobody holding its hand. That is a story, not a measurement. Here is why it matters. Alibaba's previous flagship, Qwen three point seven Max, claimed eighty point four percent on SWE-bench Verified. When vals dot ai ran the same model on a neutral harness, it measured sixty eight point eight. That eleven point six point gap is the widest on our leaderboard. So we added Qwen three point eight Max today as an unranked row. Not as a snub. It is just the only honest thing you can do with a model that has a public price and a private capability.