head-to-head

MetricClaude Fable 5GPT-5.6 Luna
SWE-bench Verified95.0%93.0%
SWE-bench Pro80.3%
Terminal-Bench
Input $ / 1M$10
Output $ / 1M$50
Context1M
Open weightsNoNo
AccessAPI · Claude Code · Claude Cowork · claude.aiAPI
MakerAnthropicOpenAI

what do the benchmarks actually say?

On SWE-bench Verified — real, human-validated GitHub issues resolved end-to-end — Claude Fable 5 posts 95.0% against 93.0% for GPT-5.6 Luna, a 2-point gap. Verified is the closest public proxy for "can it fix a real bug in a real repo without help", which is why it anchors our ranking.

A few points either way is real but not decisive: within that band, the agent scaffolding around the model — how it retrieves files, runs tests, and retries — often matters as much as the base model. Treat the gap as a lean, not a verdict.

which is cheaper to run?

Public per-token pricing isn't confirmed for both models, so we don't print a cost comparison yet.

when to pick each

Pick Claude Fable 5 if

Mythos-class flagship for long-horizon agentic runs: the model to reach for when a task spans hours and hundreds of tool calls and has to actually finish.

Pick GPT-5.6 Luna if

The cost outlier of the leaders: within striking distance of the top scores while costing a fraction per solved task, which makes it the obvious candidate for high-volume agent runs.

how were these scores verified?

We only print a number once it's confirmed against a primary source or an independent evaluation, and each row on our leaderboard records which kind it is:

  • Claude Fable 5: Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 95.00% ±0.98. Held the top score until GPT-5.6 Sol was evaluated at 96.20% ±0.86 on the same harness — a 1.2-point gap that is inside the combined margin of error (~0.9 sigma), so the two are a statistical tie and we rank Sol first only because it scored higher. SWE-bench Pro 80.3% uses Anthropic's own scaffolding and is contested. Restored Jul 1, 2026 after a 20-day export-control suspension. Pricing $10/$50 per 1M.
  • GPT-5.6 Luna: Independent (vals.ai, eval listed Jul 17 2026, mini-swe-agent bash-only harness): SWE-bench Verified 93.00% ±1.14. Added Jul 21, 2026 — this row was missing from the board even though vals.ai had already evaluated Luna, and our Kimi K3 note referenced its 93.0% score without ever listing it; adding it moves every row below it down one rank. Treat 3rd and 4th as a tie: Kimi K3's 93.40% ±1.11 is 0.4 points higher, well inside the combined margin of error (~0.25 sigma), so the ordering between them is not significant. Like the rest of the GPT-5.6 family, OpenAI has published no SWE-bench Verified figure of its own, so we rank on the independent number per our standing rule. The striking number is cost: vals.ai measured $0.21 per test against $1.15 for GPT-5.6 Sol, $1.92 for Claude Opus 4.8 and $2.05 for Claude Fable 5, at 201s median latency. Per-token list pricing not confirmed against OpenAI's own pricing page, so inPrice/outPrice stay blank rather than estimated.

Full reviewsClaude Fable 5, decoded

Ranked on our AI Coding Leaderboard, updated 2026-07-21. Scores are confirmed against primary sources; prices are per 1M input tokens and can change.

Primary sources
  • AnthropicGENZ TECH — Claude Fable 5 returns — Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 95.00% ±0.98. Held the top score until GPT-5.6 Sol was evaluated at 96.20% ±0.86 on the same harness — a 1.2-point gap that is inside the combined margin of error (~0.9 sigma), so the two are a statistical tie and we rank Sol first only because it scored higher. SWE-bench Pro 80.3% uses Anthropic's own scaffolding and is contested. Restored Jul 1, 2026 after a 20-day export-control suspension. Pricing $10/$50 per 1M.
  • OpenAIvals.ai — SWE-bench Verified (independent) — Independent (vals.ai, eval listed Jul 17 2026, mini-swe-agent bash-only harness): SWE-bench Verified 93.00% ±1.14. Added Jul 21, 2026 — this row was missing from the board even though vals.ai had already evaluated Luna, and our Kimi K3 note referenced its 93.0% score without ever listing it; adding it moves every row below it down one rank. Treat 3rd and 4th as a tie: Kimi K3's 93.40% ±1.11 is 0.4 points higher, well inside the combined margin of error (~0.25 sigma), so the ordering between them is not significant. Like the rest of the GPT-5.6 family, OpenAI has published no SWE-bench Verified figure of its own, so we rank on the independent number per our standing rule. The striking number is cost: vals.ai measured $0.21 per test against $1.15 for GPT-5.6 Sol, $1.92 for Claude Opus 4.8 and $2.05 for Claude Fable 5, at 201s median latency. Per-token list pricing not confirmed against OpenAI's own pricing page, so inPrice/outPrice stay blank rather than estimated.
  • BenchmarkSWE-bench — the real-GitHub-issue benchmark