today's standings
| # | Model | SWE-bench Verified | SWE-bench Pro | Input | Best for |
|---|---|---|---|---|---|
| 1 | Claude Opus 5 ↗Anthropic | 97.0% | — | $5/1M | The highest independently measured coding score on the board at 97.0%, at half the price of Fable 5. Strongest on short and medium tasks, though GPT-5.6 Sol still edges it on multi-hour work.Independent (vals.ai, observed Jul 25 2026, mini-swe-agent bash-only harness): SWE-bench Verified 97.00% ±0.76, the highest score on the board and… full note → |
| 2 | GPT-5.6 Sol ↗OpenAI | 96.2% | — | $5/1M | The strongest model on long tasks: 98% on the 1-to-4-hour tier, ahead of Claude Opus 5, and the second-highest overall score.Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 96.20% ±0.86 — now the second-highest score on the board… full note → |
| 4 | Claude Fable 5 ↗Anthropic | 95.0% | 80.3% | $10/1M | Mythos-class flagship for long-horizon agentic runs: the model to reach for when a task spans hours and hundreds of tool calls and has to actually finish.Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 95.00% ±0.98. Held the top score until GPT-5.6 Sol was… full note → |
| 5 | Kimi K3 open ↗Moonshot AI | 93.4% | — | $3/1M | The highest-scoring downloadable coding model we track, and by far the cheapest way to buy a 90%+ result at $3 per 1M input. Weights shipped Jul 27, 2026 as 96 shards totalling 1.56 TB, 4-bit MXFP4 only: roughly 1,454 GiB, which fits one eight-card 192GB node. No full-precision base checkpoint was published.Independent (vals.ai, mini-swe-agent bash-only harness): SWE-bench Verified 93.40% ±1.11. Verified Jul 18, 2026 — vals.ai had not run K3 at our Jul… full note → |
| 6 | GPT-5.6 Luna ↗OpenAI | 93.0% | — | $0.20/1M | The cost outlier of the leaders: within striking distance of the top scores while costing a fraction per solved task, and after an 80% price cut on Jul 30, 2026 it is now the cheapest ranked model on the board per token.Independent (vals.ai, eval listed Jul 17 2026, mini-swe-agent bash-only harness): SWE-bench Verified 93.00% ±1.14. Added Jul 21, 2026 — this row was… full note → |
| 8 | Claude Opus 4.8 ↗Anthropic | 88.6% | 69.2% | $5/1M | The hardest agentic refactors and long, autonomous multi-file tasks where every point of accuracy saves a human review cycle.Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 88.6% ±1.42. Corrected Jul 17, 2026: we previously printed… full note → |
| 9 | Grok 4.5 ↗SpaceXAI (xAI) | 86.6% | 64.7% | $2/1M | The best value at the top of the board: it solves a SWE-bench Verified task for about $2.31 of input, less than half what the two models above it cost, and it is the fastest of the leaders.Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 86.6% ±1.52. Verified Jul 17, 2026 — it launched Jul 8 with… full note → |
| 3 | Grok 4.6 ↗SpaceXAI (xAI) | 95.6% | — | $2/1M | A post-training refresh of Grok 4.5 that lands in the top group on the neutral harness, at a third the price of the models around it.Independent (vals.ai, benchmark updated 2026-08-12, mini-swe-agent bash-only harness): SWE-bench Verified 95.60% ±0.92. Entered ranked 2026-08-13… full note → |
| 10 | Muse Spark 1.2 ↗Meta | 86.6% | — | $1.25/1M | Bounded, well-specified changes at low cost, especially if you accept the contributor tier and let Meta train on your traffic.Independent (vals.ai, Aug 6 2026, mini-swe-agent bash-only harness): SWE-bench Verified 86.6%. Added Aug 7, 2026, the day after Meta launched it… full note → |
| 12 | Claude Sonnet 5 ↗Anthropic | 85.2% | 63.2% | $2/1M | The best closed-model value — near-Opus scores at ~2.5× less, and the default daily driver for most developers.Vendor-reported (Anthropic), on Anthropic's own scaffold. Independent comparison: vals.ai's bash-only harness measures Sonnet 5 at 79.6% ±1.80, 5.6… full note → |
| 13 | GLM 5.2 open ↗Z.ai (Zhipu AI) | 82.8% | — | $1.40/1M | The best open-weight coder on this board that you can actually download today: an independently measured 82.8%, MIT-licensed, and roughly a third the input price of the closed models above it.Independent (vals.ai, Jul 22 2026, mini-swe-agent bash-only harness): SWE-bench Verified 82.8% ±1.69, 9th of 75 systems. Added Jul 27, 2026 after a… full note → |
| 14 | GPT-5.5 ↗OpenAI | 82.6% | 58.6% | $5/1M | OpenAI's strongest agentic coder, with the deepest tooling and ecosystem breadth of the closed labs.Verified score from vals.ai independent eval; Pro is OpenAI-reported (rivals flag possible memorization on Pro). Price: OpenAI list $5/$30 per 1M… full note → |
| 15 | Muse Spark 1.1 ↗Meta | 82.0% | — | $1.25/1M | Meta's first paid model, and a genuine value pick: a top-10 verified score for $1.52 per solved task, cheaper per result than every model ranked above it.Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 82.0% ±1.72. Verified Jul 17, 2026; Meta published no… full note → |
| 16 | DeepSeek V4 Pro open ↗DeepSeek | 80.6% | 55.4% | $0.435/1M | The cheapest frontier-class coder — top open-weights score at ~11× less than Opus. Best pick when cost or self-hosting rules.Vendor-reported: DeepSeek's own model card, Pro-Max mode (SWE-bench Verified 80.6%, SWE-bench Pro 55.4%). Updated Aug 12, 2026: DeepSeek shipped the… full note → |
| 17 | Gemini 3.1 Pro ↗Google DeepMind | 80.6% | 54.2% | $2/1M | Google's strongest coding model today, with deep Workspace/Cloud integration. (A 3.5 Pro is expected but not shipped.)Vendor-reported (DeepMind) pass rate. No independent eval of this exact model; vals.ai has run Gemini 3.1 Pro Preview (02/26) at 78.8%, a preview… full note → |
| 18 | MiniMax M3 open ↗MiniMax | 80.5% | 59.0% | $0.60/1M | Open weights with 1M context, multimodal input and computer use — beats GPT-5.5 on SWE-bench Pro at 5–10% of the cost.Vendor-reported at launch (Jun 1, 2026). Independent comparison: vals.ai's bash-only harness measures MiniMax-M3 at 75.0% ±1.94, 5.5 points lower, so… full note → |
| 19 | Qwen3.7 Max ↗Alibaba | 80.4% | 60.6% | $1.25/1M | The best non-Claude score on the hardest benchmark — 60.6% SWE-bench Pro — built for long-horizon coding agents.Vendor-reported (May 20, 2026), and the widest vendor-versus-independent gap on this board: vals.ai's bash-only harness measures Qwen 3.7 Max at… full note → |
| 20 | Kimi K2.6 open ↗Moonshot AI | 80.2% | 58.6% | $0.95/1M | A top-three open coder whose 58.6% SWE-bench Pro beats several closed flagships.Vendor-reported (10-run average on Moonshot's own SWE-agent harness). Independent comparison: vals.ai's bash-only harness measures Kimi K2.6 at 76.2%… full note → |
| 21 | Gemini 3.6 Flash ↗Google DeepMind | 79.6% | — | $1.50/1M | Google's new general workhorse: it now outscores 3.5 Flash on an independent coding eval while emitting 17% fewer output tokens.Independent (vals.ai, mini-swe-agent bash-only harness): SWE-bench Verified 79.60% ±1.80. Verified Jul 23, 2026 — vals.ai had not run it at our Jul… full note → |
| Gemini 3.7 FlashGoogle DeepMind | — | — | $0.75/1M | Google's new volume tier, launched at half the price of 3.6 Flash. No independent coding score exists yet, so nothing to rank on.Released Aug 13, 2026, three weeks after Gemini 3.6 Flash. Google published NO SWE-bench Verified score, exactly as with 3.6 Flash, so there is… | |
| 22 | Composer 2.5 ↗Cursor | 79.6% | — | $0.50/1M | Cheap, fast in-editor coding if you already pay for Cursor. There is no way to use it anywhere else, which is the whole catch.Independent (vals.ai, Jul 22 2026, mini-swe-agent bash-only harness): SWE-bench Verified 79.6% ±1.80, 13th of 75 systems. Added Jul 27, 2026 from a… full note → |
| 23 | Gemini 3.5 Flash ↗Google DeepMind | 78.8% | — | $1.50/1M | Frontier-ish coding at Flash speed and price, with computer use built in as a native tool.vals.ai independent eval. See our decode of its native computer-use tool. Price: Google list $1.50/$9 per 1M (cached input $0.15). |
| 24 | GPT-5.6 Terra ↗OpenAI | 75.2% | — | $2/1M | Hard to justify on coding. Terra costs half of Sol but gives up 21 points of SWE-bench Verified, and Luna scores far higher for less money.Independent (vals.ai, Jul 22 2026, mini-swe-agent bash-only harness): SWE-bench Verified 75.2% ±1.93, 32nd of 75 systems. Added Jul 27, 2026 from a… full note → |
| 25 | Gemini 3.5 Flash-Lite ↗Google DeepMind | 75.0% | 54.2% | $0.30/1M | The cheapest ranked model on the board: 75.0% independently measured at $0.30 per 1M input, roughly a fifth the price of the Flash tier above it.Independent (vals.ai, mini-swe-agent bash-only harness): SWE-bench Verified 75.00% ±1.94. Verified Jul 23, 2026 — vals.ai had not run it at our Jul… full note → |
| Muse Glimmer openMeta | — | — | Free/1M | A 30B dense agentic model small enough to run always-on inside a 24GB consumer GPU, distilled from Muse Spark. Nothing to rank on yet.Released Aug 10, 2026 with weights on Hugging Face under a plain Apache 2.0 license, Meta's first genuinely permissive open-weight release in the… | |
| 11 | Qwen3.8-Max ↗Alibaba | 85.6% | — | $2.00/1M | Alibaba's 2.4T-parameter flagship, pitched at long-horizon autonomous coding. Independently scored, and slow: the highest latency of any ranked model here.Independent (vals.ai, benchmark updated 2026-08-08, mini-swe-agent bash-only harness): SWE-bench Verified 85.6% ± 1.57. Announced Aug 3, 2026 as "a… full note → |
| Laguna S 2.1 openPoolside | — | 59.4% | $0.10/1M | Frontier-adjacent agentic coding at small-model cost: a 118B mixture-of-experts with only ~8B active per token, open weights, and INT4/GGUF builds that run on local hardware.Vendor-reported (Poolside, Jul 21 2026): Terminal-Bench 2.1 70.2% with thinking enabled (60.4% without), SWE-bench Pro 59.4%, SWE-bench Multilingual… | |
| Qwen3.8-27B openAlibaba | — | 61.7% | — | A dense 27B with a vision encoder in the same stack, sized for one accelerator, and the highest published SWE-bench Pro figure of any open-weight model on this board.Vendor-reported (Alibaba, Aug 14 2026 model card): SWE-bench Pro 61.7%, Terminal-Bench 2.1 73.0, GPQA Diamond 89.2, LiveCodeBench v6 90.3, plus… | |
| LongCat-2.0 openMeituan | — | 59.5% | $0.30/1M | A 1.6T Mixture-of-Experts agentic coder trained entirely on domestic Chinese chips; led OpenRouter usage in stealth as "Owl Alpha".Vendor-reported (Meituan, Jun 30 2026): SWE-bench Pro 59.5%, Terminal-Bench 70.8%. No SWE-bench Verified score or independent eval published yet, so… | |
| Cohere North Mini Code openCohere | — | — | Free/1M | Private, self-hosted agentic coding for enterprises that cannot send source to a cloud API; runs on a single H100.Open-weight ~30B mixture-of-experts coder, free to use, runs on a single H100 with 256K context. Cohere published no SWE-bench Verified/Pro or… | |
| KAT-Coder-Pro V2.5Kwaipilot (Kuaishou) | — | 65.2% | $0.74/1M | Long-horizon agentic coding on a budget: a 72B-active MoE tuned for tool use, with a cheaper Air tier at $0.15/$0.60 on the same 256K context.Vendor-reported (KAT-Coder-V2.5 technical report, arXiv 2607.05471): SWE-bench Pro 65.2%, second only to Opus 4.8 at 69.2%, plus a best-in-test… | |
| 7 | DeepSeek V4 Flash open ↗DeepSeek | 88.8% | — | $0.14/1M | The cheapest model making a near-frontier agentic claim: 284B MoE with 13B active, 1M context, and native Responses API plus Codex support at $0.28 per 1M output.Independent (vals.ai, Aug 5 2026, mini-swe-agent bash-only harness): SWE-bench Verified 88.8%, for the DeepSeek-V4-Flash-0731 build, matching the… full note → |
| GLM-5.3Z.ai (Zhipu AI) | — | — | — | A post-training-only refresh of GLM 5.2 sold on agentic-coding token efficiency and a chart-topping vulnerability-discovery score, API-only until the weights clear safety review.Vendor-reported (Z.ai release post, Aug 14 2026). Z.ai published NO SWE-bench Verified figure for this model, so it stays unranked pending… |
how we rank
We rank by SWE-bench Verified (500 real, human-validated GitHub issues resolved end-to-end), tiebroken by the harder SWE-bench Pro. A score is only printed once confirmed against an independent evaluation or the maker's primary source — and every row states which kind it is. Where both exist, we print both: one as the ranked score, the other in that row's note. We would rather show you the gap than ask you to trust our pick. Our independent reference is vals.ai, which runs every model itself through the same minimal bash-only harness (mini-swe-agent), so the models are compared on equal footing. That matters more than it sounds: SWE-bench scores a model and its scaffolding together, and vendors report using their own tuned scaffolds. Against vals.ai's neutral harness, the vendor claims on this board run 2.6 to 11.6 points optimistic. So rows marked vendor-reported are best-case numbers and are not strictly comparable to the independent ones — where we know the independent figure, we print it in the row's note. That is a deliberate choice and you should know we made it: on four rows (Claude Sonnet 5, MiniMax M3, Qwen3.7 Max, Kimi K2.6) an independent score for that exact model exists and is lower, and we still rank on the maker's published figure because it is the number that model is sold and quoted on. We disclose the independent one beside it instead of quietly restating the board on a single evaluator's harness choice. The honest consequence: positions that straddle the two regimes are approximate. Qwen3.7 Max is the sharpest case — it sits at #16 on Alibaba's 80.4%, and on the neutral harness its 68.8% would put it far down the table. Note that llm-stats, which we previously miscredited as an independent tracker, labels its own SWE-bench Verified table "Verified: 0 / Self-reported: 104"; it aggregates vendor claims. Models still being checked are marked “verifying” and shown without a number rather than estimated. Prices are per 1M input tokens on the standard API tier and can change — always confirm current pricing with the provider.
our picks
The top independently measured score on SWE-bench Verified at 97.0%, and it costs half what Fable 5 does. Treat the lead over GPT-5.6 Sol and Fable 5 as a tie, but there is no reason to pay more for the same tier.
85.2% SWE-bench Verified at $2/1M — near-frontier coding at a rounding-error price. The default for most work.
82.8% Verified measured independently, MIT weights you can actually download, $1.40/1M. Kimi K3 scores higher and shipped its weights Jul 27, 2026 under a non-OSI license. For the absolute cheapest open option, DeepSeek V4 Pro is $0.435/1M at a vendor-reported 80.6%.
Best on the tasks that actually run long: 98% of the 1-to-4-hour SWE-bench tier, ahead of both Claude Opus 5 (90%) and Fable 5 (93%), which is what matters for unattended repo-wide work.
It's the default model for free claude.ai users — frontier-class coding at no cost for everyday tasks.
Its 60.6% on SWE-bench Pro is the best non-Claude score on the benchmark that's hardest to game.
compare head-to-head
how the field got here
- 2021GitHub Copilot preview Autocomplete-in-the-editor goes mainstream.
- 2023ChatGPT + GPT-4, then Cursor Chat-based coding and the first AI-native editor arrive.
- Aug 2024SWE-bench Verified launches A human-validated benchmark of real GitHub issues sets an honest bar.
- Oct 2024Claude 3.5 Sonnet hits ~49% Agents begin resolving real issues, not just snippets.
- 2025Terminal agents Claude Code and Codex CLI move AI out of the editor into the whole repo.
- Apr–Jun 2026Open weights close the gap DeepSeek V4, Kimi K2.6 and peers cluster at ~80% Verified — for pennies.
- 2026Verified saturates in the mid-80s SWE-bench Pro and Terminal-Bench become the real differentiators.
- Jul 2026Claude Fable 5 returns Restored after a 20-day export-control suspension; retakes SWE-bench Pro at 80.3%.
- Jun 30 2026Meituan open-sources LongCat-2.0 A 1.6T MoE coder trained entirely on domestic Chinese chips; vendor-reported 59.5% SWE-bench Pro, awaiting independent eval.
- Jul 25 2026Claude Opus 5 takes #1 vals.ai measures 97.00% ±0.76 on the neutral harness, one day after launch. A statistical tie with GPT-5.6 Sol and Claude Fable 5.
- BenchmarkSWE-bench — the real-GitHub-issue benchmark & leaderboard
- Benchmarkvals.ai — SWE-bench Verified — independent third-party evaluations
- BenchmarkSWE-bench Pro (Scale) — the harder long-horizon leaderboard
- PaperSWE-bench (arXiv) — how the benchmark is constructed
- PressGENZ TECH — Claude Fable 5 returns — restoration & SWE-bench Pro lead
- MakerAnthropic — news — Claude Opus 4.8 / Sonnet 5 model cards & pricing
- MakerOpenAI — GPT-5.5 — GPT-5.5 release post
- MakerMoonshot AI — Kimi K2.6 — open weights & reported scores
- MakerGoogle DeepMind — Gemini — Gemini model pages
- PressVentureBeat — MiniMax-M3 debut — M3 launch scores & pricing
- BenchmarkW&B ml-news — Qwen3.7-Max scores — Qwen3.7-Max benchmark table
- PressGENZ TECH — Meituan LongCat-2.0 — 1.6T MoE coder on Chinese chips (vendor-reported)
cite & embed
Free to cite and embed with attribution. Raw data: JSON · CSV. Embedding this live table adds it to your site and credits GENZ TECH.
GENZ TECH. (2026). AI Coding Leaderboard. https://genztech.blog/ai-coding-leaderboard/<iframe src="https://genztech.blog/ai-coding-leaderboard/embed/" width="100%" height="470" loading="lazy" title="AI Coding Leaderboard by GENZ TECH" style="border:1px solid #26282b;border-radius:10px;max-width:560px"></iframe>