specs at a glance

Leaderboard rank#10 of 24
SWE-bench Verified85.6%
SWE-bench Pro
Terminal-Bench
Input price / 1M$2.00
Output price / 1M$6.00
Context window1M
Open weightsNo
AccessAPI (qwen3.8-max-preview) · Qoder · QoderWork · open weights promised
MakerAlibaba

how good is Qwen3.8-Max at coding?

Qwen3.8-Max sits at #10 of 24 ranked models, posting 85.6% on SWE-bench Verified — 11.4 points behind #1 Claude Opus 5.

Score provenance: Independent (vals.ai, benchmark updated 2026-08-08, mini-swe-agent bash-only harness): SWE-bench Verified 85.6% ± 1.57. Announced Aug 3, 2026 as "a new bar for coding and cowork" at 2.4T parameters. Alibaba published NO benchmark table with the launch: no SWE-bench Verified, no SWE-bench Pro, no Terminal-Bench, no methodology, so it sat unranked here from Aug 3 until this independent score appeared. That makes it one of the few rows on this board whose only number is independent, with no vendor claim to disclose against. Two caveats worth reading with the rank. It is effectively tied with Muse Spark 1.2 and Grok 4.5 above it (86.6% each; pooled SEM on a 1.0-point gap is about 2.2, so the gap is not significant) and with Claude Sonnet 5 below. And it is by far the slowest ranked model vals.ai measured, 2502s per test against 199s for Grok 4.5, roughly 12x, at $1.13 per test. For calibration, Qwen3.7 Max carries this board's widest vendor-versus-independent gap, an 80.4% claim against 68.8% measured. Price: $2.00/$6.00 per 1M, implicit caching $0.25/1M. See /p/qwen3-8-max-2-4t-launch-no-benchmarks/.

what does Qwen3.8-Max cost?

$2.00 per 1M input tokens and $6.00 per 1M output — #15 cheapest of the 24 models we track. Coding workloads are output-heavy, so weight the output rate when budgeting. Run your own volume through the AI API cost calculator for a monthly estimate.

where can you use it?

Available via API (qwen3.8-max-preview) · Qoder · QoderWork · open weights promised. As a proprietary model, you're on the maker's infrastructure and release schedule.

Full storyQwen3.8-Max lands at 2.4T with no benchmark table

Ranked on our AI Coding Leaderboard — scores confirmed against primary sources only, updated 2026-08-10.