head-to-head
| Metric | Claude Fable 5 | GPT-5.6 Luna |
|---|---|---|
| SWE-bench Verified | 95.0% | 93.0% |
| SWE-bench Pro | 80.3% | — |
| Terminal-Bench | — | — |
| Input $ / 1M | $10 | — |
| Output $ / 1M | $50 | — |
| Context | 1M | — |
| Open weights | No | No |
| Access | API · Claude Code · Claude Cowork · claude.ai | API |
| Maker | Anthropic | OpenAI |
what do the benchmarks actually say?
On SWE-bench Verified — real, human-validated GitHub issues resolved end-to-end — Claude Fable 5 posts 95.0% against 93.0% for GPT-5.6 Luna, a 2-point gap. Verified is the closest public proxy for "can it fix a real bug in a real repo without help", which is why it anchors our ranking.
A few points either way is real but not decisive: within that band, the agent scaffolding around the model — how it retrieves files, runs tests, and retries — often matters as much as the base model. Treat the gap as a lean, not a verdict.
which is cheaper to run?
Public per-token pricing isn't confirmed for both models, so we don't print a cost comparison yet.
when to pick each
Mythos-class flagship for long-horizon agentic runs: the model to reach for when a task spans hours and hundreds of tool calls and has to actually finish.
The cost outlier of the leaders: within striking distance of the top scores while costing a fraction per solved task, which makes it the obvious candidate for high-volume agent runs.
how were these scores verified?
We only print a number once it's confirmed against a primary source or an independent evaluation, and each row on our leaderboard records which kind it is:
- Claude Fable 5: Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 95.00% ±0.98. Held the top score until GPT-5.6 Sol was evaluated at 96.20% ±0.86 on the same harness — a 1.2-point gap that is inside the combined margin of error (~0.9 sigma), so the two are a statistical tie and we rank Sol first only because it scored higher. SWE-bench Pro 80.3% uses Anthropic's own scaffolding and is contested. Restored Jul 1, 2026 after a 20-day export-control suspension. Pricing $10/$50 per 1M.
- GPT-5.6 Luna: Independent (vals.ai, eval listed Jul 17 2026, mini-swe-agent bash-only harness): SWE-bench Verified 93.00% ±1.14. Added Jul 21, 2026 — this row was missing from the board even though vals.ai had already evaluated Luna, and our Kimi K3 note referenced its 93.0% score without ever listing it; adding it moves every row below it down one rank. Treat 3rd and 4th as a tie: Kimi K3's 93.40% ±1.11 is 0.4 points higher, well inside the combined margin of error (~0.25 sigma), so the ordering between them is not significant. Like the rest of the GPT-5.6 family, OpenAI has published no SWE-bench Verified figure of its own, so we rank on the independent number per our standing rule. The striking number is cost: vals.ai measured $0.21 per test against $1.15 for GPT-5.6 Sol, $1.92 for Claude Opus 4.8 and $2.05 for Claude Fable 5, at 201s median latency. Per-token list pricing not confirmed against OpenAI's own pricing page, so inPrice/outPrice stay blank rather than estimated.
Full reviewsClaude Fable 5, decoded
Ranked on our AI Coding Leaderboard, updated 2026-07-21. Scores are confirmed against primary sources; prices are per 1M input tokens and can change.
- AnthropicGENZ TECH — Claude Fable 5 returns — Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 95.00% ±0.98. Held the top score until GPT-5.6 Sol was evaluated at 96.20% ±0.86 on the same harness — a 1.2-point gap that is inside the combined margin of error (~0.9 sigma), so the two are a statistical tie and we rank Sol first only because it scored higher. SWE-bench Pro 80.3% uses Anthropic's own scaffolding and is contested. Restored Jul 1, 2026 after a 20-day export-control suspension. Pricing $10/$50 per 1M.
- OpenAIvals.ai — SWE-bench Verified (independent) — Independent (vals.ai, eval listed Jul 17 2026, mini-swe-agent bash-only harness): SWE-bench Verified 93.00% ±1.14. Added Jul 21, 2026 — this row was missing from the board even though vals.ai had already evaluated Luna, and our Kimi K3 note referenced its 93.0% score without ever listing it; adding it moves every row below it down one rank. Treat 3rd and 4th as a tie: Kimi K3's 93.40% ±1.11 is 0.4 points higher, well inside the combined margin of error (~0.25 sigma), so the ordering between them is not significant. Like the rest of the GPT-5.6 family, OpenAI has published no SWE-bench Verified figure of its own, so we rank on the independent number per our standing rule. The striking number is cost: vals.ai measured $0.21 per test against $1.15 for GPT-5.6 Sol, $1.92 for Claude Opus 4.8 and $2.05 for Claude Fable 5, at 201s median latency. Per-token list pricing not confirmed against OpenAI's own pricing page, so inPrice/outPrice stay blank rather than estimated.
- BenchmarkSWE-bench — the real-GitHub-issue benchmark