SWE-bench Verified scores of leading AI coding modelsHorizontal bars comparing the SWE-bench Verified score of each verified frontier coding model, against the late-2024 frontier baseline of about 49 percent.SWE-BENCH VERIFIED (% RESOLVED)2024 ≈ 49%Claude Opus 597% GPT-5.6 Sol96.2% Grok 4.695.6% Claude Fable 595% Kimi K393.4% GPT-5.6 Luna93% DeepSeek V4 Flash88.8% Claude Opus 4.888.6% Grok 4.586.6% Muse Spark 1.286.6% Qwen3.8-Max85.6% Claude Sonnet 585.2% GLM 5.282.8% GPT-5.582.6% Muse Spark 1.182% DeepSeek V4 Pro80.6% Gemini 3.1 Pro80.6% MiniMax M380.5% Qwen3.7 Max80.4% Kimi K2.680.2% Gemini 3.6 Flash79.6% Composer 2.579.6% Gemini 3.5 Flash78.8% GPT-5.6 Terra75.2% Gemini 3.5 Flash-Lite75%</> genztech.blog
Fig 1 · benchmark SWE-bench Verified — the share of real, human-validated GitHub issues an AI resolves end-to-end. Only models whose scores we have confirmed are charted; models still being verified are left off rather than estimated. Sources: vals.ai independent evaluations where available, otherwise maker reports — each row states which. Vendor-reported scores use the vendor's own scaffold and run high.

today's standings

#ModelSWE-bench VerifiedSWE-bench ProInputBest for
1 Claude Opus 5 Anthropic 97.0% $5/1M The highest independently measured coding score on the board at 97.0%, at half the price of Fable 5. Strongest on short and medium tasks, though GPT-5.6 Sol still edges it on multi-hour work.Independent (vals.ai, observed Jul 25 2026, mini-swe-agent bash-only harness): SWE-bench Verified 97.00% ±0.76, the highest score on the board and… full note →
2 GPT-5.6 Sol OpenAI 96.2% $5/1M The strongest model on long tasks: 98% on the 1-to-4-hour tier, ahead of Claude Opus 5, and the second-highest overall score.Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 96.20% ±0.86 — now the second-highest score on the board… full note →
4 Claude Fable 5 Anthropic 95.0% 80.3% $10/1M Mythos-class flagship for long-horizon agentic runs: the model to reach for when a task spans hours and hundreds of tool calls and has to actually finish.Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 95.00% ±0.98. Held the top score until GPT-5.6 Sol was… full note →
5 Kimi K3 open Moonshot AI 93.4% $3/1M The highest-scoring downloadable coding model we track, and by far the cheapest way to buy a 90%+ result at $3 per 1M input. Weights shipped Jul 27, 2026 as 96 shards totalling 1.56 TB, 4-bit MXFP4 only: roughly 1,454 GiB, which fits one eight-card 192GB node. No full-precision base checkpoint was published.Independent (vals.ai, mini-swe-agent bash-only harness): SWE-bench Verified 93.40% ±1.11. Verified Jul 18, 2026 — vals.ai had not run K3 at our Jul… full note →
6 GPT-5.6 Luna OpenAI 93.0% $0.20/1M The cost outlier of the leaders: within striking distance of the top scores while costing a fraction per solved task, and after an 80% price cut on Jul 30, 2026 it is now the cheapest ranked model on the board per token.Independent (vals.ai, eval listed Jul 17 2026, mini-swe-agent bash-only harness): SWE-bench Verified 93.00% ±1.14. Added Jul 21, 2026 — this row was… full note →
8 Claude Opus 4.8 Anthropic 88.6% 69.2% $5/1M The hardest agentic refactors and long, autonomous multi-file tasks where every point of accuracy saves a human review cycle.Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 88.6% ±1.42. Corrected Jul 17, 2026: we previously printed… full note →
9 Grok 4.5 SpaceXAI (xAI) 86.6% 64.7% $2/1M The best value at the top of the board: it solves a SWE-bench Verified task for about $2.31 of input, less than half what the two models above it cost, and it is the fastest of the leaders.Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 86.6% ±1.52. Verified Jul 17, 2026 — it launched Jul 8 with… full note →
3 Grok 4.6 SpaceXAI (xAI) 95.6% $2/1M A post-training refresh of Grok 4.5 that lands in the top group on the neutral harness, at a third the price of the models around it.Independent (vals.ai, benchmark updated 2026-08-12, mini-swe-agent bash-only harness): SWE-bench Verified 95.60% ±0.92. Entered ranked 2026-08-13… full note →
10 Muse Spark 1.2 Meta 86.6% $1.25/1M Bounded, well-specified changes at low cost, especially if you accept the contributor tier and let Meta train on your traffic.Independent (vals.ai, Aug 6 2026, mini-swe-agent bash-only harness): SWE-bench Verified 86.6%. Added Aug 7, 2026, the day after Meta launched it… full note →
12 Claude Sonnet 5 Anthropic 85.2% 63.2% $2/1M The best closed-model value — near-Opus scores at ~2.5× less, and the default daily driver for most developers.Vendor-reported (Anthropic), on Anthropic's own scaffold. Independent comparison: vals.ai's bash-only harness measures Sonnet 5 at 79.6% ±1.80, 5.6… full note →
13 GLM 5.2 open Z.ai (Zhipu AI) 82.8% $1.40/1M The best open-weight coder on this board that you can actually download today: an independently measured 82.8%, MIT-licensed, and roughly a third the input price of the closed models above it.Independent (vals.ai, Jul 22 2026, mini-swe-agent bash-only harness): SWE-bench Verified 82.8% ±1.69, 9th of 75 systems. Added Jul 27, 2026 after a… full note →
14 GPT-5.5 OpenAI 82.6% 58.6% $5/1M OpenAI's strongest agentic coder, with the deepest tooling and ecosystem breadth of the closed labs.Verified score from vals.ai independent eval; Pro is OpenAI-reported (rivals flag possible memorization on Pro). Price: OpenAI list $5/$30 per 1M… full note →
15 Muse Spark 1.1 Meta 82.0% $1.25/1M Meta's first paid model, and a genuine value pick: a top-10 verified score for $1.52 per solved task, cheaper per result than every model ranked above it.Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 82.0% ±1.72. Verified Jul 17, 2026; Meta published no… full note →
16 DeepSeek V4 Pro open DeepSeek 80.6% 55.4% $0.435/1M The cheapest frontier-class coder — top open-weights score at ~11× less than Opus. Best pick when cost or self-hosting rules.Vendor-reported: DeepSeek's own model card, Pro-Max mode (SWE-bench Verified 80.6%, SWE-bench Pro 55.4%). Updated Aug 12, 2026: DeepSeek shipped the… full note →
17 Gemini 3.1 Pro Google DeepMind 80.6% 54.2% $2/1M Google's strongest coding model today, with deep Workspace/Cloud integration. (A 3.5 Pro is expected but not shipped.)Vendor-reported (DeepMind) pass rate. No independent eval of this exact model; vals.ai has run Gemini 3.1 Pro Preview (02/26) at 78.8%, a preview… full note →
18 MiniMax M3 open MiniMax 80.5% 59.0% $0.60/1M Open weights with 1M context, multimodal input and computer use — beats GPT-5.5 on SWE-bench Pro at 5–10% of the cost.Vendor-reported at launch (Jun 1, 2026). Independent comparison: vals.ai's bash-only harness measures MiniMax-M3 at 75.0% ±1.94, 5.5 points lower, so… full note →
19 Qwen3.7 Max Alibaba 80.4% 60.6% $1.25/1M The best non-Claude score on the hardest benchmark — 60.6% SWE-bench Pro — built for long-horizon coding agents.Vendor-reported (May 20, 2026), and the widest vendor-versus-independent gap on this board: vals.ai's bash-only harness measures Qwen 3.7 Max at… full note →
20 Kimi K2.6 open Moonshot AI 80.2% 58.6% $0.95/1M A top-three open coder whose 58.6% SWE-bench Pro beats several closed flagships.Vendor-reported (10-run average on Moonshot's own SWE-agent harness). Independent comparison: vals.ai's bash-only harness measures Kimi K2.6 at 76.2%… full note →
21 Gemini 3.6 Flash Google DeepMind 79.6% $1.50/1M Google's new general workhorse: it now outscores 3.5 Flash on an independent coding eval while emitting 17% fewer output tokens.Independent (vals.ai, mini-swe-agent bash-only harness): SWE-bench Verified 79.60% ±1.80. Verified Jul 23, 2026 — vals.ai had not run it at our Jul… full note →
verifying Gemini 3.7 FlashGoogle DeepMind $0.75/1M Google's new volume tier, launched at half the price of 3.6 Flash. No independent coding score exists yet, so nothing to rank on.Released Aug 13, 2026, three weeks after Gemini 3.6 Flash. Google published NO SWE-bench Verified score, exactly as with 3.6 Flash, so there is…
22 Composer 2.5 Cursor 79.6% $0.50/1M Cheap, fast in-editor coding if you already pay for Cursor. There is no way to use it anywhere else, which is the whole catch.Independent (vals.ai, Jul 22 2026, mini-swe-agent bash-only harness): SWE-bench Verified 79.6% ±1.80, 13th of 75 systems. Added Jul 27, 2026 from a… full note →
23 Gemini 3.5 Flash Google DeepMind 78.8% $1.50/1M Frontier-ish coding at Flash speed and price, with computer use built in as a native tool.vals.ai independent eval. See our decode of its native computer-use tool. Price: Google list $1.50/$9 per 1M (cached input $0.15).
24 GPT-5.6 Terra OpenAI 75.2% $2/1M Hard to justify on coding. Terra costs half of Sol but gives up 21 points of SWE-bench Verified, and Luna scores far higher for less money.Independent (vals.ai, Jul 22 2026, mini-swe-agent bash-only harness): SWE-bench Verified 75.2% ±1.93, 32nd of 75 systems. Added Jul 27, 2026 from a… full note →
25 Gemini 3.5 Flash-Lite Google DeepMind 75.0% 54.2% $0.30/1M The cheapest ranked model on the board: 75.0% independently measured at $0.30 per 1M input, roughly a fifth the price of the Flash tier above it.Independent (vals.ai, mini-swe-agent bash-only harness): SWE-bench Verified 75.00% ±1.94. Verified Jul 23, 2026 — vals.ai had not run it at our Jul… full note →
verifying Muse Glimmer openMeta Free/1M A 30B dense agentic model small enough to run always-on inside a 24GB consumer GPU, distilled from Muse Spark. Nothing to rank on yet.Released Aug 10, 2026 with weights on Hugging Face under a plain Apache 2.0 license, Meta's first genuinely permissive open-weight release in the…
11 Qwen3.8-Max Alibaba 85.6% $2.00/1M Alibaba's 2.4T-parameter flagship, pitched at long-horizon autonomous coding. Independently scored, and slow: the highest latency of any ranked model here.Independent (vals.ai, benchmark updated 2026-08-08, mini-swe-agent bash-only harness): SWE-bench Verified 85.6% ± 1.57. Announced Aug 3, 2026 as "a… full note →
verifying Laguna S 2.1 openPoolside 59.4% $0.10/1M Frontier-adjacent agentic coding at small-model cost: a 118B mixture-of-experts with only ~8B active per token, open weights, and INT4/GGUF builds that run on local hardware.Vendor-reported (Poolside, Jul 21 2026): Terminal-Bench 2.1 70.2% with thinking enabled (60.4% without), SWE-bench Pro 59.4%, SWE-bench Multilingual…
verifying Qwen3.8-27B openAlibaba 61.7% A dense 27B with a vision encoder in the same stack, sized for one accelerator, and the highest published SWE-bench Pro figure of any open-weight model on this board.Vendor-reported (Alibaba, Aug 14 2026 model card): SWE-bench Pro 61.7%, Terminal-Bench 2.1 73.0, GPQA Diamond 89.2, LiveCodeBench v6 90.3, plus…
verifying LongCat-2.0 openMeituan 59.5% $0.30/1M A 1.6T Mixture-of-Experts agentic coder trained entirely on domestic Chinese chips; led OpenRouter usage in stealth as "Owl Alpha".Vendor-reported (Meituan, Jun 30 2026): SWE-bench Pro 59.5%, Terminal-Bench 70.8%. No SWE-bench Verified score or independent eval published yet, so…
verifying Cohere North Mini Code openCohere Free/1M Private, self-hosted agentic coding for enterprises that cannot send source to a cloud API; runs on a single H100.Open-weight ~30B mixture-of-experts coder, free to use, runs on a single H100 with 256K context. Cohere published no SWE-bench Verified/Pro or…
verifying KAT-Coder-Pro V2.5Kwaipilot (Kuaishou) 65.2% $0.74/1M Long-horizon agentic coding on a budget: a 72B-active MoE tuned for tool use, with a cheaper Air tier at $0.15/$0.60 on the same 256K context.Vendor-reported (KAT-Coder-V2.5 technical report, arXiv 2607.05471): SWE-bench Pro 65.2%, second only to Opus 4.8 at 69.2%, plus a best-in-test…
7 DeepSeek V4 Flash open DeepSeek 88.8% $0.14/1M The cheapest model making a near-frontier agentic claim: 284B MoE with 13B active, 1M context, and native Responses API plus Codex support at $0.28 per 1M output.Independent (vals.ai, Aug 5 2026, mini-swe-agent bash-only harness): SWE-bench Verified 88.8%, for the DeepSeek-V4-Flash-0731 build, matching the… full note →
verifying GLM-5.3Z.ai (Zhipu AI) A post-training-only refresh of GLM 5.2 sold on agentic-coding token efficiency and a chart-topping vulnerability-discovery score, API-only until the weights clear safety review.Vendor-reported (Z.ai release post, Aug 14 2026). Z.ai published NO SWE-bench Verified figure for this model, so it stays unranked pending…

how we rank

We rank by SWE-bench Verified (500 real, human-validated GitHub issues resolved end-to-end), tiebroken by the harder SWE-bench Pro. A score is only printed once confirmed against an independent evaluation or the maker's primary source — and every row states which kind it is. Where both exist, we print both: one as the ranked score, the other in that row's note. We would rather show you the gap than ask you to trust our pick. Our independent reference is vals.ai, which runs every model itself through the same minimal bash-only harness (mini-swe-agent), so the models are compared on equal footing. That matters more than it sounds: SWE-bench scores a model and its scaffolding together, and vendors report using their own tuned scaffolds. Against vals.ai's neutral harness, the vendor claims on this board run 2.6 to 11.6 points optimistic. So rows marked vendor-reported are best-case numbers and are not strictly comparable to the independent ones — where we know the independent figure, we print it in the row's note. That is a deliberate choice and you should know we made it: on four rows (Claude Sonnet 5, MiniMax M3, Qwen3.7 Max, Kimi K2.6) an independent score for that exact model exists and is lower, and we still rank on the maker's published figure because it is the number that model is sold and quoted on. We disclose the independent one beside it instead of quietly restating the board on a single evaluator's harness choice. The honest consequence: positions that straddle the two regimes are approximate. Qwen3.7 Max is the sharpest case — it sits at #16 on Alibaba's 80.4%, and on the neutral harness its 68.8% would put it far down the table. Note that llm-stats, which we previously miscredited as an independent tracker, labels its own SWE-bench Verified table "Verified: 0 / Self-reported: 104"; it aggregates vendor claims. Models still being checked are marked “verifying” and shown without a number rather than estimated. Prices are per 1M input tokens on the standard API tier and can change — always confirm current pricing with the provider.

our picks

Best overallClaude Opus 5

The top independently measured score on SWE-bench Verified at 97.0%, and it costs half what Fable 5 does. Treat the lead over GPT-5.6 Sol and Fable 5 as a tie, but there is no reason to pay more for the same tier.

Best value (closed)Claude Sonnet 5

85.2% SWE-bench Verified at $2/1M — near-frontier coding at a rounding-error price. The default for most work.

Best open weightsGLM 5.2

82.8% Verified measured independently, MIT weights you can actually download, $1.40/1M. Kimi K3 scores higher and shipped its weights Jul 27, 2026 under a non-OSI license. For the absolute cheapest open option, DeepSeek V4 Pro is $0.435/1M at a vendor-reported 80.6%.

Best for autonomous agentsGPT-5.6 Sol

Best on the tasks that actually run long: 98% of the 1-to-4-hour SWE-bench tier, ahead of both Claude Opus 5 (90%) and Fable 5 (93%), which is what matters for unattended repo-wide work.

Best free optionClaude Sonnet 5

It's the default model for free claude.ai users — frontier-class coding at no cost for everyday tasks.

Hardest-tasks dark horseQwen3.7 Max

Its 60.6% on SWE-bench Pro is the best non-Claude score on the benchmark that's hardest to game.

compare head-to-head

how the field got here

Top SWE-bench Verified score over time, Jul 2 to Aug 13 2026A line chart of the best confirmed SWE-bench Verified score on this leaderboard at each date the ranked data changed. It runs from 86 percent (Claude Opus 4.8) on Jul 2 to 97 percent (Claude Opus 5) on Aug 13, with 3 changes of leader. Points mark the dates we recorded a change; the axis is to scale, so gaps are periods with no movement.TOP SWE-BENCH VERIFIED SCORE OVER TIMEeach point = a date our ranked data changed · axis to scale84%92%99%Jul 2 Jul 3 Jul 17 Jul 18 Jul 19 Jul 21 Jul 23 Jul 25 Jul 27 Aug 1 Aug 6 Aug 7 Aug 10 Aug 13~86%Claude Opus 4.8 95.0%Claude Fable 5 96.2%GPT-5.6 Sol 96.2%GPT-5.6 Sol 96.2%GPT-5.6 Sol 96.2%GPT-5.6 Sol 96.2%GPT-5.6 Sol 97.0%Claude Opus 5 97.0%Claude Opus 5 97.0%Claude Opus 5 97.0%Claude Opus 5 97.0%Claude Opus 5 97.0%Claude Opus 5 97.0%Claude Opus 5</> genztech.blog
Fig 2 · history The frontier since we started tracking on Jul 2 2026: +11 points in 42 days, across 3 changes of leader. We plot one point per date the ranked data actually moved — not per edit — and the axis is to scale, so a flat stretch means nothing changed rather than that we stopped looking. Full series: JSON · CSV, CC BY 4.0.
  1. 2021GitHub Copilot preview Autocomplete-in-the-editor goes mainstream.
  2. 2023ChatGPT + GPT-4, then Cursor Chat-based coding and the first AI-native editor arrive.
  3. Aug 2024SWE-bench Verified launches A human-validated benchmark of real GitHub issues sets an honest bar.
  4. Oct 2024Claude 3.5 Sonnet hits ~49% Agents begin resolving real issues, not just snippets.
  5. 2025Terminal agents Claude Code and Codex CLI move AI out of the editor into the whole repo.
  6. Apr–Jun 2026Open weights close the gap DeepSeek V4, Kimi K2.6 and peers cluster at ~80% Verified — for pennies.
  7. 2026Verified saturates in the mid-80s SWE-bench Pro and Terminal-Bench become the real differentiators.
  8. Jul 2026Claude Fable 5 returns Restored after a 20-day export-control suspension; retakes SWE-bench Pro at 80.3%.
  9. Jun 30 2026Meituan open-sources LongCat-2.0 A 1.6T MoE coder trained entirely on domestic Chinese chips; vendor-reported 59.5% SWE-bench Pro, awaiting independent eval.
  10. Jul 25 2026Claude Opus 5 takes #1 vals.ai measures 97.00% ±0.76 on the neutral harness, one day after launch. A statistical tie with GPT-5.6 Sol and Claude Fable 5.
Primary sources

cite & embed

Free to cite and embed with attribution. Raw data: JSON · CSV. Embedding this live table adds it to your site and credits GENZ TECH.

CiteGENZ TECH. (2026). AI Coding Leaderboard. https://genztech.blog/ai-coding-leaderboard/
Embed<iframe src="https://genztech.blog/ai-coding-leaderboard/embed/" width="100%" height="470" loading="lazy" title="AI Coding Leaderboard by GENZ TECH" style="border:1px solid #26282b;border-radius:10px;max-width:560px"></iframe>