specs at a glance
| Leaderboard rank | #11 of 28 |
|---|---|
| SWE-bench Verified | 86.6% |
| SWE-bench Pro | 64.7% |
| Terminal-Bench | 83.3% (TB2.1) |
| Input price / 1M | $2 |
| Output price / 1M | $6 |
| Context window | — |
| Open weights | No |
| Access | API · SpaceXAI console · Grok Build · Cursor (all plans); not in EU yet |
| Maker | SpaceXAI (xAI) |
how good is Grok 4.5 at coding?
Grok 4.5 sits at #11 of 28 ranked models, posting 86.6% on SWE-bench Verified — 10.4 points behind #1 Claude Opus 5. On the harder SWE-bench Pro it scores 64.7%. Terminal-Bench (agentic terminal work): 83.3% (TB2.1).
Score provenance: Independent (vals.ai, Jul 14 2026, mini-swe-agent bash-only harness): SWE-bench Verified 86.6% ±1.52. Verified Jul 17, 2026 — it launched Jul 8 with no Verified score and we had it unranked as "Opus-class, unproven"; the independent number now backs the Opus-class claim, landing it 2 points under Claude Opus 4.8 at 40% of the input price. On our price-per-solved-task metric that is ~$2.31 against $5.64 for Opus 4.8 and $10.53 for Fable 5, the cheapest of any model scoring above 85%. It is also quick: 199.6s mean latency per task versus 566.9s for Opus 4.8. SWE-bench Pro 64.7% and Terminal-Bench 2.1 83.3% remain SpaceXAI-reported. Priced $2/$6 per 1M. Still not available in the EU.
what does Grok 4.5 cost?
$2 per 1M input tokens and $6 per 1M output — #17 cheapest of the 26 models we track. Coding workloads are output-heavy, so weight the output rate when budgeting. Run your own volume through the AI API cost calculator for a monthly estimate.
where can you use it?
Available via API · SpaceXAI console · Grok Build · Cursor (all plans); not in EU yet. As a proprietary model, you're on the maker's infrastructure and release schedule.
Full storyGrok 4.5 lands: Opus-class claims, cheaper, unproven
- SpaceXAI (xAI)vals.ai — SWE-bench Verified (independent)
Ranked on our AI Coding Leaderboard — scores confirmed against primary sources only, updated 2026-08-20.