SpaceXAI released Grok 4.6 late Wednesday, roughly an hour ago as this publishes, and the number it wants you to take away is 61: an Artificial Analysis Intelligence Index score that ties GPT-5.6 Sol Max and sits one point under Claude Fable 5 Max. That part is real. The eval table underneath it tells a more specific story, and it is not the one the headline implies. Grok 4.6 is now genuinely frontier-class at knowledge work and professional tasks. At driving a terminal it is still the weakest of the four models xAI put on its own chart.
The model is a post-training release, not a new foundation. Same lineage as Grok 4.5, a longer supplemental training run, regenerated supervised fine-tuning trajectories, and agentic reinforcement learning across kernel optimization, web development and computer-aided design. Pricing did not move: $2 per million input tokens and $6 per million output, with a fast variant at double that.
RelatedQwen3.8-Max lands at 2.4T with no benchmark table
- It ties the frontier on the composite and loses on the components that matter to agents. AA Intelligence Index 61, level with GPT-5.6 Sol Max. Terminal-Bench v3.0 is 26% against Sol's 34.6%.
- Where it actually leads is knowledge work. Grok 4.6 posts the best score in xAI's table on GDPVal-AA (1753), AA-Briefcase (1577) and Harvey LAB (15.8%), the three least code-shaped evals on the list.
- The price held flat across a generation. $2/$6 is unchanged from Grok 4.5 and roughly a third of what the models beating it on DeepSWE cost.
- Nothing here is independently verified yet. Every number in this piece is vendor-published, including the rival scores, which xAI assembled from other labs' system cards.
What did SpaceXAI actually ship?
Grok 4.6 went live in Cursor and Grok Build first, and is available through the SpaceXAI API plus OpenRouter, Vercel and Cloudflare. xAI is running 2x included usage inside Grok Build and Cursor for the first week, which is the usual land-grab move on launch day and worth using if you were going to evaluate it anyway.
The training description is unusually candid about what this is. xAI says it used Grok 4.5 itself to regenerate the SFT trajectories across reasoning efforts, agent harnesses and domains, then filtered bad traces with model-based checks. That is distillation-flavoured self-improvement on top of an existing base, and it explains the shape of the gains: big jumps on the tasks the RL environments targeted, small movement everywhere else.
Cursor, which had the model early, described it as strong at turning a broad product idea into a working first version and noted stronger first passes on visual and interactive projects than Grok 4.5 gave them. xAI's own framing matches: it says longer trajectories showed the model doing more self-testing and verification before moving on.
Where does Grok 4.6 win, and where does it lose?
Six of the ten evals xAI published are on a percentage scale, which makes them directly comparable. On three of those, Grok 4.6 comes out on top of GPT-5.6 Sol Max, and every one of those wins is inside a rounding error: CursorBench v3.2 by 2.7 points over Sol but 0.6 behind Fable 5, FrontierCode by 0.7 over Sol, APEX-Agents by 0.8 over Sol. Call those ties.
The losses are not ties. DeepSWE v1.1 puts Grok 4.6 at 65.9% against Sol's 73%, a 7.1-point gap. Terminal-Bench v3.0 is worse: 26% against 34.6% for Sol and 34.1% for Fable 5. On a benchmark where nobody is doing well, Grok 4.6 is solving roughly three quarters as many tasks as the leaders.
Then there is the other half of the table, the index-scale evals, and the pattern flips completely. GDPVal-AA v2, which measures economically valuable knowledge work rather than code, has Grok 4.6 at 1753, ahead of both Fable 5 Max (1741) and Sol Max (1728). AA-Briefcase: 1577 against 1574 and 1502. Harvey LAB, a legal-reasoning eval run by Vals: 15.8% against Fable 5's 11.3% and Sol's 2.5%. Grok 4.6 does not merely win that last one, it wins it by more than 6x over Sol.
| Eval | Grok 4.6 High | Grok 4.5 High | GPT-5.6 Sol Max | Fable 5 Max |
|---|---|---|---|---|
| AA Intelligence Index | 61 | 56 | 61 | 62 |
| GDPVal-AA v2 | 1753 | 1526 | 1728 | 1741 |
| CursorBench v3.2 | 69.9% | 66.7% | 67.2% | 70.5% |
| DeepSWE v1.1 | 65.9% | 54% | 73% | 70% |
| FrontierCode v1.1 (ext) | 61.3% | 56.6% | 60.6% | 63.6% |
| APEX-Agents | 57.5% | 47.1% | 56.7% | 59.2% |
| Terminal-Bench v3.0 | 26% | 15.7% | 34.6% | 34.1% |
| APEX-SWE | 56.4% | 53.6% | not published | 58.8% |
| AA-Briefcase | 1577 | 1313 | 1502 | 1574 |
| Harvey LAB (Vals) | 15.8% | 12.9% | 2.5% | 11.3% |
Why are the Terminal-Bench numbers so low?
Because this is v3.0, and it is a different benchmark from the one everyone quoted three weeks ago. Grok 4.5 shipped with a SpaceXAI-reported Terminal-Bench 2.1 score of 83.3%. The same model scores 15.7% on v3.0. Nothing regressed; the test got much harder, and the whole field collapsed with it. Sol and Fable 5, both at roughly 34%, are the current ceiling.
This matters more than a version bump usually would, because Terminal-Bench is the eval closest to what people actually deploy coding models to do: hand them a shell and a goal and leave. A 26% pass rate means three out of four attempts fail. Every lab is in that boat right now, which is a useful correction to the "AI does software engineering" framing that a 90-percent SWE-bench Verified number invites.
Read Fig 2 alongside Fig 1 and the strategy is legible. The three evals with double-digit gains from 4.5 to 4.6, DeepSWE (+11.9), APEX-Agents (+10.4) and Terminal-Bench (+10.3), are exactly the three where xAI still trails. The RL environments xAI describes, kernel optimization and web development among them, were pointed at the gap. It closed some of it in one generation. It did not close all of it.
RelatedGPT-5.6 Sol's #1 coding score is basically a tie
Is this a coding model or a knowledge-work model?
xAI is selling it as a coding agent, and that is where the distribution is, in Cursor and Grok Build. But the evidence in its own table says the differentiated capability is elsewhere. A model that beats GPT-5.6 Sol on legal reasoning by 13 points and on a general economic-work benchmark by 25 index points, while losing the terminal by 8, is not primarily a coding model. It is a generalist that codes competently and reasons about professional documents unusually well.
If you are picking a model this week, that distinction is the whole decision. For agentic terminal work with long autonomous runs, Sol and Fable 5 remain the better bets on the published evidence. For research, analysis, document-heavy workflows and first-pass application scaffolding, Grok 4.6 at $2 per million input is a serious value argument, and Cursor's read that it produces strong initial versions of visual and interactive projects lines up with that.
- Jul 8, 2026Grok 4.5 ships with no SWE-bench Verified score We ranked it unranked and called it unproven at the time.
- Jul 17, 2026vals.ai measures Grok 4.5 at 86.6% on a neutral harness It entered our leaderboard at #8, the cheapest model above 85%.
- Aug 12, 2026Grok 4.6 launches at the same $2/$6 pricing Live in Cursor, Grok Build, the SpaceXAI API, OpenRouter, Vercel and Cloudflare.
- PendingAn independent SWE-bench Verified run for Grok 4.6 Until one exists it sits unranked on our board with no score, the same way 4.5 did.
What it means for the market
The signal for investors is not the benchmark, it is the price tag that did not move. xAI held Grok 4.6 at $2/$6 across a full generation of capability gains, which puts a model at parity on the AA composite roughly a third below what the comparably-scoring alternatives charge. That is deliberate pressure on per-token margins at the top of the market, and it lands the same month Alibaba's 2.4T Qwen3.8-Max arrived at $2/$6 as well.
Public-market exposure here is indirect, since xAI now sits inside SpaceX and is not separately traded. Watch the distribution layer instead: Cloudflare (NET) and the inference-routing platforms carry Grok 4.6 from day one, and each additional frontier model that clears the quality bar at half the price makes routing-and-arbitrage businesses more valuable and single-vendor API lock-in less so. Nvidia demand is unaffected either way; a post-training release consumes compute regardless of whose logo is on it. This is analysis, not investment advice.
Our take
Treat the 61 as marketing and the table as the product. xAI deserves credit for publishing ten evals including the ones it loses, which is more disclosure than most launch posts carry, and for footnoting that the rival scores are the best self-reported or publicly available numbers rather than its own runs. That footnote is also the reason to hold judgment: a comparison table assembled by one vendor from other vendors' best published results is the most favourable possible construction of the field, and every score in it, including Grok's, comes from the lab that benefits.
Our leaderboard adds Grok 4.6 today with no score and no rank, which is the honest position 90 minutes after launch. Grok 4.5 sat in exactly that state for nine days before vals.ai ran it on a neutral bash-only harness and the 86.6% came back, at which point the Opus-class claim held up. The same test is what 4.6 needs. If it lands anywhere near 4.5's number while carrying these knowledge-work scores, the value argument at $2 per million gets hard to ignore.
- An independent SWE-bench Verified run. vals.ai took nine days on Grok 4.5. Until that number exists, 4.6 is a vendor claim with good provenance and nothing more.
- Whether Terminal-Bench v3.0 stays the ceiling. Nobody is above 35%. The first lab to clear 50% there has a genuinely different product, not a better score.
- Whether the knowledge-work lead survives contact. GDPVal-AA and Harvey LAB are the interesting results here, and they are the ones fewest people will replicate.
- Grok 4.7. xAI's own cadence puts 4.5 and 4.6 five weeks apart. Another point release before October would confirm this is a rapid post-training loop rather than a roadmap.
- OfficialSpaceXAI: Introducing Grok 4.6 — the launch post and the full ten-eval comparison table, published Aug 12, 2026.
- PartnerCursor: Grok 4.6 — early-access notes, pricing and the fast-variant multiplier.
- Benchmarkvals.ai SWE-bench Verified — the independent bash-only harness that scored Grok 4.5 at 86.6%.
- DataGENZ TECH AI Coding Leaderboard — our ranked board; Grok 4.6 enters unranked pending an independent score.
Original analysis by GenZTech, built from SpaceXAI's published eval table and our own leaderboard data. Launch details confirmed against x.ai/news/grok-4-6.
