SpaceXAI released Grok 4.6 late Wednesday, roughly an hour ago as this publishes, and the number it wants you to take away is 61: an Artificial Analysis Intelligence Index score that ties GPT-5.6 Sol Max and sits one point under Claude Fable 5 Max. That part is real. The eval table underneath it tells a more specific story, and it is not the one the headline implies. Grok 4.6 is now genuinely frontier-class at knowledge work and professional tasks. At driving a terminal it is still the weakest of the four models xAI put on its own chart.

The model is a post-training release, not a new foundation. Same lineage as Grok 4.5, a longer supplemental training run, regenerated supervised fine-tuning trajectories, and agentic reinforcement learning across kernel optimization, web development and computer-aided design. Pricing did not move: $2 per million input tokens and $6 per million output, with a fast variant at double that.

RelatedQwen3.8-Max lands at 2.4T with no benchmark table

  • It ties the frontier on the composite and loses on the components that matter to agents. AA Intelligence Index 61, level with GPT-5.6 Sol Max. Terminal-Bench v3.0 is 26% against Sol's 34.6%.
  • Where it actually leads is knowledge work. Grok 4.6 posts the best score in xAI's table on GDPVal-AA (1753), AA-Briefcase (1577) and Harvey LAB (15.8%), the three least code-shaped evals on the list.
  • The price held flat across a generation. $2/$6 is unchanged from Grok 4.5 and roughly a third of what the models beating it on DeepSWE cost.
  • Nothing here is independently verified yet. Every number in this piece is vendor-published, including the rival scores, which xAI assembled from other labs' system cards.

What did SpaceXAI actually ship?

Grok 4.6 went live in Cursor and Grok Build first, and is available through the SpaceXAI API plus OpenRouter, Vercel and Cloudflare. xAI is running 2x included usage inside Grok Build and Cursor for the first week, which is the usual land-grab move on launch day and worth using if you were going to evaluate it anyway.

The training description is unusually candid about what this is. xAI says it used Grok 4.5 itself to regenerate the SFT trajectories across reasoning efforts, agent harnesses and domains, then filtered bad traces with model-based checks. That is distillation-flavoured self-improvement on top of an existing base, and it explains the shape of the gains: big jumps on the tasks the RL environments targeted, small movement everywhere else.

Cursor, which had the model early, described it as strong at turning a broad product idea into a working first version and noted stronger first passes on visual and interactive projects than Grok 4.5 gave them. xAI's own framing matches: it says longer trajectories showed the model doing more self-testing and verification before moving on.

Grok 4.6 against GPT-5.6 Sol Max and Fable 5 Max on six percentage-scale agentic evals Grouped bar chart. Grok 4.6 leads on CursorBench, FrontierCode and APEX-Agents by under one point, and trails badly on DeepSWE and Terminal-Bench version 3.0. PERCENTAGE-SCALE EVALS / XAI-PUBLISHED TABLE / AUG 12 2026 Grok 4.6 High GPT-5.6 Sol Max Fable 5 Max 020406080 n/a 69.965.961.357.556.426.0 CursorBenchv3.2 DeepSWEv1.1 FrontierCodev1.1 ext APEXAgents APEXSWE Terminal-Benchv3.0 genztech.blog
Fig 1 · benchmark The three evals Grok 4.6 wins here it wins by 0.7 points or less. The two it loses, DeepSWE and Terminal-Bench v3.0, it loses by 7.1 and 8.6 points. All figures from SpaceXAI's own launch table; rival scores are the best self-reported or public numbers xAI could find, not xAI's own runs.

Where does Grok 4.6 win, and where does it lose?

Six of the ten evals xAI published are on a percentage scale, which makes them directly comparable. On three of those, Grok 4.6 comes out on top of GPT-5.6 Sol Max, and every one of those wins is inside a rounding error: CursorBench v3.2 by 2.7 points over Sol but 0.6 behind Fable 5, FrontierCode by 0.7 over Sol, APEX-Agents by 0.8 over Sol. Call those ties.

The losses are not ties. DeepSWE v1.1 puts Grok 4.6 at 65.9% against Sol's 73%, a 7.1-point gap. Terminal-Bench v3.0 is worse: 26% against 34.6% for Sol and 34.1% for Fable 5. On a benchmark where nobody is doing well, Grok 4.6 is solving roughly three quarters as many tasks as the leaders.

Then there is the other half of the table, the index-scale evals, and the pattern flips completely. GDPVal-AA v2, which measures economically valuable knowledge work rather than code, has Grok 4.6 at 1753, ahead of both Fable 5 Max (1741) and Sol Max (1728). AA-Briefcase: 1577 against 1574 and 1502. Harvey LAB, a legal-reasoning eval run by Vals: 15.8% against Fable 5's 11.3% and Sol's 2.5%. Grok 4.6 does not merely win that last one, it wins it by more than 6x over Sol.

EvalGrok 4.6 HighGrok 4.5 HighGPT-5.6 Sol MaxFable 5 Max
AA Intelligence Index61566162
GDPVal-AA v21753152617281741
CursorBench v3.269.9%66.7%67.2%70.5%
DeepSWE v1.165.9%54%73%70%
FrontierCode v1.1 (ext)61.3%56.6%60.6%63.6%
APEX-Agents57.5%47.1%56.7%59.2%
Terminal-Bench v3.026%15.7%34.6%34.1%
APEX-SWE56.4%53.6%not published58.8%
AA-Briefcase1577131315021574
Harvey LAB (Vals)15.8%12.9%2.5%11.3%

Why are the Terminal-Bench numbers so low?

Because this is v3.0, and it is a different benchmark from the one everyone quoted three weeks ago. Grok 4.5 shipped with a SpaceXAI-reported Terminal-Bench 2.1 score of 83.3%. The same model scores 15.7% on v3.0. Nothing regressed; the test got much harder, and the whole field collapsed with it. Sol and Fable 5, both at roughly 34%, are the current ceiling.

This matters more than a version bump usually would, because Terminal-Bench is the eval closest to what people actually deploy coding models to do: hand them a shell and a goal and leave. A 26% pass rate means three out of four attempts fail. Every lab is in that boat right now, which is a useful correction to the "AI does software engineering" framing that a 90-percent SWE-bench Verified number invites.

Percentage-point gain from Grok 4.5 to Grok 4.6 on six evals Horizontal bar chart. The largest gains land on DeepSWE, APEX-Agents and Terminal-Bench, the same three evals where Grok 4.6 sits furthest behind its rivals. GROK 4.5 HIGH TO GROK 4.6 HIGH / POINTS GAINED The generation moved most where the model was weakest. DeepSWE v1.1APEX-AgentsTerminal-Bench v3.0FrontierCode v1.1CursorBench v3.2APEX-SWE +11.9+10.4+10.3+4.7+3.2+2.8 SOURCE: X.AI/NEWS/GROK-4-6 EVAL TABLE genztech.blog
Fig 2 · benchmark Grok 4.6's three biggest jumps over Grok 4.5 are on DeepSWE, APEX-Agents and Terminal-Bench. Those are also the three evals where it still ranks last or near-last in xAI's own table, which is what a post-training push into a known weakness looks like.

Read Fig 2 alongside Fig 1 and the strategy is legible. The three evals with double-digit gains from 4.5 to 4.6, DeepSWE (+11.9), APEX-Agents (+10.4) and Terminal-Bench (+10.3), are exactly the three where xAI still trails. The RL environments xAI describes, kernel optimization and web development among them, were pointed at the gap. It closed some of it in one generation. It did not close all of it.

RelatedGPT-5.6 Sol's #1 coding score is basically a tie

Is this a coding model or a knowledge-work model?

xAI is selling it as a coding agent, and that is where the distribution is, in Cursor and Grok Build. But the evidence in its own table says the differentiated capability is elsewhere. A model that beats GPT-5.6 Sol on legal reasoning by 13 points and on a general economic-work benchmark by 25 index points, while losing the terminal by 8, is not primarily a coding model. It is a generalist that codes competently and reasons about professional documents unusually well.

If you are picking a model this week, that distinction is the whole decision. For agentic terminal work with long autonomous runs, Sol and Fable 5 remain the better bets on the published evidence. For research, analysis, document-heavy workflows and first-pass application scaffolding, Grok 4.6 at $2 per million input is a serious value argument, and Cursor's read that it produces strong initial versions of visual and interactive projects lines up with that.

  1. Jul 8, 2026Grok 4.5 ships with no SWE-bench Verified score We ranked it unranked and called it unproven at the time.
  2. Jul 17, 2026vals.ai measures Grok 4.5 at 86.6% on a neutral harness It entered our leaderboard at #8, the cheapest model above 85%.
  3. Aug 12, 2026Grok 4.6 launches at the same $2/$6 pricing Live in Cursor, Grok Build, the SpaceXAI API, OpenRouter, Vercel and Cloudflare.
  4. PendingAn independent SWE-bench Verified run for Grok 4.6 Until one exists it sits unranked on our board with no score, the same way 4.5 did.

What it means for the market

The signal for investors is not the benchmark, it is the price tag that did not move. xAI held Grok 4.6 at $2/$6 across a full generation of capability gains, which puts a model at parity on the AA composite roughly a third below what the comparably-scoring alternatives charge. That is deliberate pressure on per-token margins at the top of the market, and it lands the same month Alibaba's 2.4T Qwen3.8-Max arrived at $2/$6 as well.

Public-market exposure here is indirect, since xAI now sits inside SpaceX and is not separately traded. Watch the distribution layer instead: Cloudflare (NET) and the inference-routing platforms carry Grok 4.6 from day one, and each additional frontier model that clears the quality bar at half the price makes routing-and-arbitrage businesses more valuable and single-vendor API lock-in less so. Nvidia demand is unaffected either way; a post-training release consumes compute regardless of whose logo is on it. This is analysis, not investment advice.

Our take

Treat the 61 as marketing and the table as the product. xAI deserves credit for publishing ten evals including the ones it loses, which is more disclosure than most launch posts carry, and for footnoting that the rival scores are the best self-reported or publicly available numbers rather than its own runs. That footnote is also the reason to hold judgment: a comparison table assembled by one vendor from other vendors' best published results is the most favourable possible construction of the field, and every score in it, including Grok's, comes from the lab that benefits.

Our leaderboard adds Grok 4.6 today with no score and no rank, which is the honest position 90 minutes after launch. Grok 4.5 sat in exactly that state for nine days before vals.ai ran it on a neutral bash-only harness and the 86.6% came back, at which point the Opus-class claim held up. The same test is what 4.6 needs. If it lands anywhere near 4.5's number while carrying these knowledge-work scores, the value argument at $2 per million gets hard to ignore.

What to watch · next 2-4 weeks
  • An independent SWE-bench Verified run. vals.ai took nine days on Grok 4.5. Until that number exists, 4.6 is a vendor claim with good provenance and nothing more.
  • Whether Terminal-Bench v3.0 stays the ceiling. Nobody is above 35%. The first lab to clear 50% there has a genuinely different product, not a better score.
  • Whether the knowledge-work lead survives contact. GDPVal-AA and Harvey LAB are the interesting results here, and they are the ones fewest people will replicate.
  • Grok 4.7. xAI's own cadence puts 4.5 and 4.6 five weeks apart. Another point release before October would confirm this is a rapid post-training loop rather than a roadmap.
Primary sources

Original analysis by GenZTech, built from SpaceXAI's published eval table and our own leaderboard data. Launch details confirmed against x.ai/news/grok-4-6.