SpaceXAI shipped Grok 4.7 on the morning of September 21, calling it its most powerful model yet for coding and knowledge work. The release lands about nine days after Elon Musk's own mid-September target, and it arrives priced identically to its predecessor while leaning on a new benchmark suite that makes head-to-head comparisons harder than usual.
- Grok 4.7 is live now through the Grok API, Grok Build, Cursor, and unnamed third-party platforms, priced at $2 per million input tokens and $6 per million output tokens, the same headline rate as Grok 4.6.
- SpaceXAI says the model leads on price-performance on CursorBench 4.0 at 46.3%, a new benchmark version that isn't directly comparable to Grok 4.6's 69.9% score on the older CursorBench 3.2.
- The company also published scores on three professional-domain evals: 64% on an electrical engineering benchmark, 56.7% on clinical reasoning, and 19.6% on a legal-work test.
- No SWE-bench Verified score has been published or independently measured yet, and the industry's main independent scoreboard, run by vals.ai, has been archived since September 5 over score saturation.
What did SpaceXAI actually announce?
The company's own release notes call Grok 4.7 its most powerful model for coding and knowledge work, built on what it describes only as "a new, larger base model" that went through extended reinforcement learning aimed at complex, multi-hour problems rather than single-turn answers. SpaceXAI also says the model got native integration with the Grok Bot harness, which it credits with lifting general-conversation and knowledge-recall performance rather than just coding. A faster-serving variant is available at double the price for double the speed, the same trade SpaceXAI already offers on Grok 4.6.
RelatedGrok 4.6 Ties GPT-5.6 Sol but Loses the Terminal
Why did the release slip from September 11 to the 21st?
Musk first teased Grok 4.7 on September 2, targeting a September 11 or 12 launch and claiming, ahead of any public benchmark, that it would beat every existing model. That date came and went with no official release, and as of September 18 there was still no confirmed ship date. SpaceXAI hasn't explained the nine-day slip in its announcement post, which is itself a pattern: Grok 4.5 and Grok 4.6 both shipped without vendor-measured SWE-bench Verified scores attached, with the harder coding numbers arriving later or from outside evaluators. A short, unexplained delay on a model that was publicly promised as imminent is a minor thing on its own, but it's the same shape of gap between what gets announced and what ships with receipts attached.
What's the actual mechanism behind the multi-hour claim?
The interesting technical claim here isn't the model size, it's the training target. Most frontier labs have spent 2026 optimizing for benchmark tasks that resolve in minutes: a single bug fix, a single file edit, a single terminal command. SpaceXAI says Grok 4.7's reinforcement learning explicitly targeted longer-running, self-verifying work, the kind of task where a coding agent has to check its own output against a spec across many steps without a human re-reading every intermediate result. That's a real gap in current agentic coding: models are good at producing a plausible next step and much weaker at noticing when step 40 contradicts step 12. Whether Grok 4.7 actually closes that gap isn't something a single vendor-published benchmark can settle, since self-verification is exactly the kind of capability that's easy to claim and hard to test without an independent, long-horizon benchmark, which doesn't widely exist yet.
How does it score, and why can't you compare it directly to Grok 4.6?
SpaceXAI's headline number, 46.3% on CursorBench 4.0, sounds like a step down from Grok 4.6's 69.9% on CursorBench 3.2, but the two aren't the same test. Cursor's benchmark maker revises the suite periodically to keep pace with what models can already solve, the same reason Terminal-Bench jumped from a version where Grok 4.5 scored 83.3% to a harder revision where the same model scores 15.7%. A new, harder CursorBench 4.0 makes 46.3% a genuinely different claim than 69.9% on the old version, and SpaceXAI's own framing, "leads in price-performance," is doing real work: it's a claim about value per dollar among models tested on the new suite, not a claim of outright highest score. The three professional-domain numbers, 64% electrical engineering, 56.7% clinical reasoning, 19.6% legal work, are presented with no comparison baseline against rival models at all, so they're only useful as a snapshot of where SpaceXAI thinks the model is strong or weak, not as a ranking.
Who actually has to care about this today?
Developers already on Cursor get Grok 4.7 as a model option without doing anything, since Cursor has been a SpaceX-owned subsidiary since the $60 billion Anysphere acquisition closed on August 14. That ownership structure matters more than usual here: SpaceXAI now controls both the model and one of the most widely used editors it ships in, which is a different incentive shape than a neutral API provider competing for developer attention. Teams already paying Grok 4.6 rates on the API see no price change, so switching is a config edit, not a budget conversation. Anyone evaluating coding models purely on a leaderboard should hold off: with no independent SWE-bench Verified score and vals.ai's board archived since September 5, there's currently no third-party number to check SpaceXAI's claims against, on this model or any other released since then.
RelatedGrok 4.5 lands: Opus-class claims, cheaper, unproven
What it means for the AI infrastructure race
SpaceX is private, so there's no ticker to move, but the signal is still worth reading. Folding Cursor into the same division as Grok Build and Grok Bot, then shipping a model tuned specifically for longer coding sessions, points at SpaceXAI building a vertically integrated coding stack rather than just selling model access. That's a different bet than OpenAI or Anthropic make when they ship a model and let third-party tools like Cursor or Cline plug into it. If SpaceXAI can make Grok 4.7 meaningfully better inside Cursor specifically, using signals or fine-tuning a pure API competitor can't access, that's a real moat. If it can't, owning the editor mostly just adds overhead. Investors and enterprise buyers evaluating SpaceX's broader AI ambitions should watch adoption inside Cursor specifically, not just API pricing, as the tell.
| Grok 4.5 | Grok 4.6 | Grok 4.7 | |
|---|---|---|---|
| Launched | Jul 8, 2026 | Aug 12, 2026 | Sep 21, 2026 |
| Pricing (in/out per 1M) | $2 / $6 | $2 / $6 | $2 / $6 |
| Independent SWE-bench Verified | 86.6% (vals.ai) | 95.6% (vals.ai) | Not yet measured |
| Headline vendor benchmark | CursorBench v3.2 66.7% | CursorBench v3.2 69.9% | CursorBench v4.0 46.3% |
| Access | API, Grok Build, Cursor | + OpenRouter, Vercel, Cloudflare | API, Grok Build, Cursor, unnamed 3rd parties |
- Jul 8, 2026Grok 4.5 launches no SWE-bench Verified score at launch
- Aug 12, 2026Grok 4.6 ships vals.ai measures 95.6% within a day
- Aug 14, 2026SpaceX closes $60B Cursor acquisition folds Anysphere into the SpaceXAI division
- Sep 2, 2026Musk teases Grok 4.7 targets a Sep 11-12 release
- Sep 5, 2026vals.ai archives its SWE-bench Verified board cites score saturation
- Sep 11-18, 2026Target date missed no official ship date confirmed
- Sep 21, 2026Grok 4.7 ships CursorBench 4.0 and three professional-domain evals published, no SWE-bench Verified
- Independent scoring. With vals.ai's board archived, watch for whichever benchmark operator picks up the slack, or whether SpaceXAI eventually publishes its own SWE-bench Verified number the way it delayed doing for 4.5 and 4.6.
- Cursor-specific gains. If Grok 4.7 pulls meaningfully ahead of Grok 4.6 specifically inside Cursor rather than through the raw API, that's the vertical-integration bet paying off.
- The multi-hour claim under real load. Long-running, self-verifying agent tasks are hard to fake in a demo but also hard to benchmark; expect early developer reports, not another vendor chart, to be the real signal.
Our take
Grok 4.7 is a real release with real availability today, not a teaser, and the multi-hour, self-verifying training focus targets a genuine weak spot in current coding agents. But SpaceXAI shipped it the same way it shipped the last two Grok models: strong on vendor-chosen benchmarks, silent on the one number, SWE-bench Verified, that this site's own leaderboard uses to compare across the whole market, and unlucky enough to launch the week after the industry's main independent scoreboard went dark. None of that means the model is weak. It means nobody outside SpaceXAI can currently say how it stacks up, and a nine-day slip past a self-announced date with no public explanation is a small but real crack in the "beats everything" framing Musk used to tease it.
- OfficialSpaceXAI, Introducing Grok 4.7 launch announcement, pricing, and benchmark claims
- ReferenceGenZTech, Grok 4.6 ties GPT-5.6 Sol but loses the terminal our prior coverage and independent SWE-bench Verified numbers
- ReferenceGenZTech, SpaceX closes $60B Cursor acquisition how Cursor became part of the SpaceXAI division
- BenchmarkGenZTech AI Coding Leaderboard independent SWE-bench Verified rankings, updated as scores are confirmed
Original analysis by GenZTech, based on SpaceXAI's official announcement and our own benchmark tracking. Read SpaceXAI's release notes.
