head-to-head

MetricClaude Opus 5Grok 4.6
SWE-bench Verified97.0%95.6%
SWE-bench Pro
Terminal-Bench
Input $ / 1M$5$2
Output $ / 1M$25$6
Context
Open weightsNoNo
AccessAPI (claude-opus-5) · Claude Code · Claude Cowork · claude.ai (Pro, Max)API · Grok Build · Cursor · OpenRouter · Vercel · Cloudflare; fast variant at 2x price
MakerAnthropicSpaceXAI (xAI)

what do the benchmarks actually say?

On SWE-bench Verified — real, human-validated GitHub issues resolved end-to-end — Claude Opus 5 posts 97.0% against 95.6% for Grok 4.6, a 1.4-point gap. Verified is the closest public proxy for "can it fix a real bug in a real repo without help", which is why it anchors our ranking.

A few points either way is real but not decisive: within that band, the agent scaffolding around the model — how it retrieves files, runs tests, and retries — often matters as much as the base model. Treat the gap as a lean, not a verdict.

which is cheaper to run?

Grok 4.6 is the cheaper model: $2 per 1M input tokens ($6 output) versus $5 ($25 output) for Claude Opus 5 — roughly 2.5× less on input. Coding workloads are output-heavy — agents write diffs, tests and retries — so weight the output rate more than the input rate when you estimate a monthly bill.

when to pick each

Pick Claude Opus 5 if

The highest independently measured coding score on the board at 97.0%, at half the price of Fable 5. Strongest on short and medium tasks, though GPT-5.6 Sol still edges it on multi-hour work.

Pick Grok 4.6 if

A post-training refresh of Grok 4.5 that lands in the top group on the neutral harness, at a third the price of the models around it.

how were these scores verified?

We only print a number once it's confirmed against a primary source or an independent evaluation, and each row on our leaderboard records which kind it is:

  • Claude Opus 5: Independent (vals.ai, observed Jul 25 2026, mini-swe-agent bash-only harness): SWE-bench Verified 97.00% ±0.76, the highest score on the board and 1st of the 75 systems vals.ai has run. Entered ranked Jul 25, 2026 after one day unranked: Anthropic published no SWE-bench Verified number at launch and still has not, so this is vals.ai's own measurement rather than a vendor claim. Read the #1 as a three-way tie, not a win — GPT-5.6 Sol is at 96.20% ±0.86 (a 0.8-point gap, ~0.7 sigma) and Claude Fable 5 at 95.00% ±0.98 (2.0 points, ~1.6 sigma), both inside the combined margin of error. We rank Opus 5 first only because it scored highest. Where the top two genuinely separate is task length, and not in Opus 5's favour: on the 1-to-4-hour tier Sol solves 98% against Opus 5's 90%, while Opus 5 leads on shorter work (98% under 15 minutes and 97% on 15-minute-to-1-hour tasks, vs 97% and 95% for Sol). Released Jul 24, 2026 at $5/$25 per 1M, the same price as Opus 4.8 and half of Fable 5. Anthropic's launch claims stay unreproducible (Frontier-Bench v0.1, CursorBench 3.2 and Zapier AutomationBench are proprietary), so this is the first externally checkable score the model has. Fast mode runs about 2.5x default speed at 2x base price.
  • Grok 4.6: Independent (vals.ai, benchmark updated 2026-08-12, mini-swe-agent bash-only harness): SWE-bench Verified 95.60% ±0.92. Entered ranked 2026-08-13, one day after launch, and it enters near the top: 4th of the 82 systems vals.ai has run, behind Claude Opus 5 (97.00% ±0.76), DeepSeek V4 Pro 0813 (96.40% ±0.83) and GPT-5.6 Sol (96.20% ±0.86). Read the gap to Claude Fable 5 below it (95.00% ±0.98) as a tie: 0.6 points against a pooled SEM of about 1.34 is well inside the margin of error, and we rank Grok 4.6 higher only because it scored higher. Same against Sol above it, a 0.6-point gap. This is the second time running that xAI shipped a Grok with no SWE-bench Verified number of its own and vals.ai supplied one: Grok 4.5 waited nine days, Grok 4.6 waited one. SpaceXAI still publishes no SWE-bench Verified or SWE-bench Pro figure, so the ranked score here is entirely vals.ai's measurement with no vendor claim to disclose against. What xAI did publish, all vendor-reported on its own table: AA Intelligence Index 61, CursorBench v3.2 69.9%, DeepSWE v1.1 65.9%, FrontierCode v1.1 Extended 61.3%, APEX-Agents 57.5%, APEX-SWE 56.4%, Terminal-Bench v3.0 26%, GDPVal-AA v2 1753, AA-Briefcase 1577, Harvey LAB (Vals) 15.8%. That table's shape still holds and is worth reading against this score: Grok 4.6 takes the best result on the three knowledge-work evals and loses the two hardest agentic-coding ones by wide margins (DeepSWE 65.9% vs Sol's 73%, Terminal-Bench v3.0 26% vs Sol's 34.6%). Terminal-Bench v3.0 is NOT comparable to the 83.3% TB2.1 figure on the Grok 4.5 row: on v3.0 Grok 4.5 scores 15.7% and the whole field tops out near 34%. On cost it is the standout of the top five, $0.78 per test against $1.29 for Opus 5, $1.15 for Sol and $2.05 for Fable 5, though it is also the slowest of them at 604s. Pricing unchanged from Grok 4.5 at $2/$6 per 1M. See /p/grok-4-6-frontier-benchmarks-terminal-bench-gap/.

Full reviewsClaude Opus 5, decodedGrok 4.6, decoded

Ranked on our AI Coding Leaderboard, updated 2026-08-13. Scores are confirmed against primary sources; prices are per 1M input tokens and can change.

Primary sources
  • Anthropicvals.ai — SWE-bench Verified (independent) — Independent (vals.ai, observed Jul 25 2026, mini-swe-agent bash-only harness): SWE-bench Verified 97.00% ±0.76, the highest score on the board and 1st of the 75 systems vals.ai has run. Entered ranked Jul 25, 2026 after one day unranked: Anthropic published no SWE-bench Verified number at launch and still has not, so this is vals.ai's own measurement rather than a vendor claim. Read the #1 as a three-way tie, not a win — GPT-5.6 Sol is at 96.20% ±0.86 (a 0.8-point gap, ~0.7 sigma) and Claude Fable 5 at 95.00% ±0.98 (2.0 points, ~1.6 sigma), both inside the combined margin of error. We rank Opus 5 first only because it scored highest. Where the top two genuinely separate is task length, and not in Opus 5's favour: on the 1-to-4-hour tier Sol solves 98% against Opus 5's 90%, while Opus 5 leads on shorter work (98% under 15 minutes and 97% on 15-minute-to-1-hour tasks, vs 97% and 95% for Sol). Released Jul 24, 2026 at $5/$25 per 1M, the same price as Opus 4.8 and half of Fable 5. Anthropic's launch claims stay unreproducible (Frontier-Bench v0.1, CursorBench 3.2 and Zapier AutomationBench are proprietary), so this is the first externally checkable score the model has. Fast mode runs about 2.5x default speed at 2x base price.
  • SpaceXAI (xAI)vals.ai — SWE-bench Verified (independent) — Independent (vals.ai, benchmark updated 2026-08-12, mini-swe-agent bash-only harness): SWE-bench Verified 95.60% ±0.92. Entered ranked 2026-08-13, one day after launch, and it enters near the top: 4th of the 82 systems vals.ai has run, behind Claude Opus 5 (97.00% ±0.76), DeepSeek V4 Pro 0813 (96.40% ±0.83) and GPT-5.6 Sol (96.20% ±0.86). Read the gap to Claude Fable 5 below it (95.00% ±0.98) as a tie: 0.6 points against a pooled SEM of about 1.34 is well inside the margin of error, and we rank Grok 4.6 higher only because it scored higher. Same against Sol above it, a 0.6-point gap. This is the second time running that xAI shipped a Grok with no SWE-bench Verified number of its own and vals.ai supplied one: Grok 4.5 waited nine days, Grok 4.6 waited one. SpaceXAI still publishes no SWE-bench Verified or SWE-bench Pro figure, so the ranked score here is entirely vals.ai's measurement with no vendor claim to disclose against. What xAI did publish, all vendor-reported on its own table: AA Intelligence Index 61, CursorBench v3.2 69.9%, DeepSWE v1.1 65.9%, FrontierCode v1.1 Extended 61.3%, APEX-Agents 57.5%, APEX-SWE 56.4%, Terminal-Bench v3.0 26%, GDPVal-AA v2 1753, AA-Briefcase 1577, Harvey LAB (Vals) 15.8%. That table's shape still holds and is worth reading against this score: Grok 4.6 takes the best result on the three knowledge-work evals and loses the two hardest agentic-coding ones by wide margins (DeepSWE 65.9% vs Sol's 73%, Terminal-Bench v3.0 26% vs Sol's 34.6%). Terminal-Bench v3.0 is NOT comparable to the 83.3% TB2.1 figure on the Grok 4.5 row: on v3.0 Grok 4.5 scores 15.7% and the whole field tops out near 34%. On cost it is the standout of the top five, $0.78 per test against $1.29 for Opus 5, $1.15 for Sol and $2.05 for Fable 5, though it is also the slowest of them at 604s. Pricing unchanged from Grok 4.5 at $2/$6 per 1M. See /p/grok-4-6-frontier-benchmarks-terminal-bench-gap/.
  • BenchmarkSWE-bench — the real-GitHub-issue benchmark