specs at a glance

Leaderboard rank#3 of 25
SWE-bench Verified95.6%
SWE-bench Pro
Terminal-Bench
Input price / 1M$2
Output price / 1M$6
Context window
Open weightsNo
AccessAPI · Grok Build · Cursor · OpenRouter · Vercel · Cloudflare; fast variant at 2x price
MakerSpaceXAI (xAI)

how good is Grok 4.6 at coding?

Grok 4.6 sits at #3 of 25 ranked models, posting 95.6% on SWE-bench Verified — 1.4 points behind #1 Claude Opus 5.

Score provenance: Independent (vals.ai, benchmark updated 2026-08-12, mini-swe-agent bash-only harness): SWE-bench Verified 95.60% ±0.92. Entered ranked 2026-08-13, one day after launch, and it enters near the top: 4th of the 82 systems vals.ai has run, behind Claude Opus 5 (97.00% ±0.76), DeepSeek V4 Pro 0813 (96.40% ±0.83) and GPT-5.6 Sol (96.20% ±0.86). Read the gap to Claude Fable 5 below it (95.00% ±0.98) as a tie: 0.6 points against a pooled SEM of about 1.34 is well inside the margin of error, and we rank Grok 4.6 higher only because it scored higher. Same against Sol above it, a 0.6-point gap. This is the second time running that xAI shipped a Grok with no SWE-bench Verified number of its own and vals.ai supplied one: Grok 4.5 waited nine days, Grok 4.6 waited one. SpaceXAI still publishes no SWE-bench Verified or SWE-bench Pro figure, so the ranked score here is entirely vals.ai's measurement with no vendor claim to disclose against. What xAI did publish, all vendor-reported on its own table: AA Intelligence Index 61, CursorBench v3.2 69.9%, DeepSWE v1.1 65.9%, FrontierCode v1.1 Extended 61.3%, APEX-Agents 57.5%, APEX-SWE 56.4%, Terminal-Bench v3.0 26%, GDPVal-AA v2 1753, AA-Briefcase 1577, Harvey LAB (Vals) 15.8%. That table's shape still holds and is worth reading against this score: Grok 4.6 takes the best result on the three knowledge-work evals and loses the two hardest agentic-coding ones by wide margins (DeepSWE 65.9% vs Sol's 73%, Terminal-Bench v3.0 26% vs Sol's 34.6%). Terminal-Bench v3.0 is NOT comparable to the 83.3% TB2.1 figure on the Grok 4.5 row: on v3.0 Grok 4.5 scores 15.7% and the whole field tops out near 34%. On cost it is the standout of the top five, $0.78 per test against $1.29 for Opus 5, $1.15 for Sol and $2.05 for Fable 5, though it is also the slowest of them at 604s. Pricing unchanged from Grok 4.5 at $2/$6 per 1M. See /p/grok-4-6-frontier-benchmarks-terminal-bench-gap/.

what does Grok 4.6 cost?

$2 per 1M input tokens and $6 per 1M output — #14 cheapest of the 25 models we track. Coding workloads are output-heavy, so weight the output rate when budgeting. Run your own volume through the AI API cost calculator for a monthly estimate.

where can you use it?

Available via API · Grok Build · Cursor · OpenRouter · Vercel · Cloudflare; fast variant at 2x price. As a proprietary model, you're on the maker's infrastructure and release schedule.

head-to-head

Full storyGrok 4.6 Ties GPT-5.6 Sol but Loses the Terminal

Ranked on our AI Coding Leaderboard — scores confirmed against primary sources only, updated 2026-08-13.