head-to-head
| Metric | GPT-5.6 Sol | GLM-5.3 |
|---|---|---|
| SWE-bench Verified | 96.2% | 95.4% |
| SWE-bench Pro | — | — |
| Terminal-Bench | 88.8% | 28.3 (TB3.0, vendor) |
| Input $ / 1M | $5 | — |
| Output $ / 1M | $30 | — |
| Context | — | 1M |
| Open weights | No | No |
| Access | API · Codex (public since Jul 9 2026) | API (glm-5.3) · GLM Coding Plan · ZCode; weights promised ~2 weeks after launch |
| Maker | OpenAI | Z.ai (Zhipu AI) |
what do the benchmarks actually say?
On SWE-bench Verified — real, human-validated GitHub issues resolved end-to-end — GPT-5.6 Sol posts 96.2% against 95.4% for GLM-5.3, a 0.8-point gap. Verified is the closest public proxy for "can it fix a real bug in a real repo without help", which is why it anchors our ranking.
A few points either way is real but not decisive: within that band, the agent scaffolding around the model — how it retrieves files, runs tests, and retries — often matters as much as the base model. Treat the gap as a lean, not a verdict.
which is cheaper to run?
Public per-token pricing isn't confirmed for both models, so we don't print a cost comparison yet.
when to pick each
The strongest model on long tasks: 98% on the 1-to-4-hour tier, ahead of Claude Opus 5, and the second-highest overall score.
The highest-scoring model on this board whose maker has promised open weights, and the cheapest route to a 95%-plus measured score at $0.34 per test.
how were these scores verified?
We only print a number once it's confirmed against a primary source or an independent evaluation, and each row on our leaderboard records which kind it is:
- GPT-5.6 Sol: Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 96.20% ±0.86 — now the second-highest score on the board. Correction, Jul 25, 2026: this row read "the top score on the board" from Jul 17 until Jul 25, when vals.ai evaluated Claude Opus 5 at 97.00% ±0.76 and took the #1 slot. The 0.8-point gap is ~0.7 sigma and not significant, so the two are a statistical tie, and Sol still leads on the longest tasks (98% vs 90% on the 1-to-4-hour tier). Verified Jul 17, 2026; it had been unranked since Jun 26 because OpenAI published no SWE-bench number of its own, and it still has not. Read the #1 with care: the 1.2-point lead over Claude Fable 5 (95.00% ±0.98) is inside the combined margin of error (~0.9 sigma, not significant), so the two are a statistical tie and we rank Sol first only because it scored higher. Where it does separate is task length — 98% on 1-4 hour tasks vs 93% for Fable 5. OpenAI's own Terminal-Bench 2.1 claim is 88.8% (Sol) / 91.9% (Sol Ultra). No SWE-bench Pro score published. Pricing $5/$30 per 1M.
- GLM-5.3: Independent (vals.ai, 2026-08-20 sweep, mini-swe-agent bash-only harness): SWE-bench Verified 95.4% ±0.94, 6th of 86 systems. Left the verifying queue on 2026-08-20, six days after launch. Z.ai published NO SWE-bench Verified figure of its own, so this rank rests entirely on the independent run, which is the cleanest kind of row on this board. Read it as tied with GPT-5.6 Terra (95.4% ±0.94), Grok 4.6 (95.6%) and Claude Fable 5 (95.0%); pooled SEM is about 1.3 points and every gap is inside it. That result is striking given the release is post-training only, on the same base model as GLM 5.2, which is ranked far below on an independent 82.8%. Vendor numbers from the Aug 14 2026 release post, run mostly inside a Claude Code 2.1.207 harness at max effort rather than a neutral one: Terminal-Bench 3.0 28.3 (up from GLM-5.2’s 4.6), DeepSWE v1.1 66.9 (from 46.2), Agents’ Last Exam 28.5 (from 23.8). Its headline claims were cyber rather than coding: CyberGym 84.5% and ExploitBench 54.4%. Marked closed because the weights are NOT out: Z.ai said roughly two weeks after launch pending safety hardening, with no license announced, so the open-weights flag flips only when they actually land. See /p/glm-5-3-cybergym-open-weights-delayed/.
Full reviewsGPT-5.6 Sol, decodedGLM-5.3, decoded
Ranked on our AI Coding Leaderboard, updated 2026-08-20. Scores are confirmed against primary sources; prices are per 1M input tokens and can change.
- OpenAIvals.ai — SWE-bench Verified (independent) — Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 96.20% ±0.86 — now the second-highest score on the board. Correction, Jul 25, 2026: this row read "the top score on the board" from Jul 17 until Jul 25, when vals.ai evaluated Claude Opus 5 at 97.00% ±0.76 and took the #1 slot. The 0.8-point gap is ~0.7 sigma and not significant, so the two are a statistical tie, and Sol still leads on the longest tasks (98% vs 90% on the 1-to-4-hour tier). Verified Jul 17, 2026; it had been unranked since Jun 26 because OpenAI published no SWE-bench number of its own, and it still has not. Read the #1 with care: the 1.2-point lead over Claude Fable 5 (95.00% ±0.98) is inside the combined margin of error (~0.9 sigma, not significant), so the two are a statistical tie and we rank Sol first only because it scored higher. Where it does separate is task length — 98% on 1-4 hour tasks vs 93% for Fable 5. OpenAI's own Terminal-Bench 2.1 claim is 88.8% (Sol) / 91.9% (Sol Ultra). No SWE-bench Pro score published. Pricing $5/$30 per 1M.
- Z.ai (Zhipu AI)vals.ai — SWE-bench Verified (independent) — Independent (vals.ai, 2026-08-20 sweep, mini-swe-agent bash-only harness): SWE-bench Verified 95.4% ±0.94, 6th of 86 systems. Left the verifying queue on 2026-08-20, six days after launch. Z.ai published NO SWE-bench Verified figure of its own, so this rank rests entirely on the independent run, which is the cleanest kind of row on this board. Read it as tied with GPT-5.6 Terra (95.4% ±0.94), Grok 4.6 (95.6%) and Claude Fable 5 (95.0%); pooled SEM is about 1.3 points and every gap is inside it. That result is striking given the release is post-training only, on the same base model as GLM 5.2, which is ranked far below on an independent 82.8%. Vendor numbers from the Aug 14 2026 release post, run mostly inside a Claude Code 2.1.207 harness at max effort rather than a neutral one: Terminal-Bench 3.0 28.3 (up from GLM-5.2’s 4.6), DeepSWE v1.1 66.9 (from 46.2), Agents’ Last Exam 28.5 (from 23.8). Its headline claims were cyber rather than coding: CyberGym 84.5% and ExploitBench 54.4%. Marked closed because the weights are NOT out: Z.ai said roughly two weeks after launch pending safety hardening, with no license announced, so the open-weights flag flips only when they actually land. See /p/glm-5-3-cybergym-open-weights-delayed/.
- BenchmarkSWE-bench — the real-GitHub-issue benchmark