Grok 4.6 is out, and the headline is not the intelligence number but the fact that the price did not move. Artificial Analysis measures it at 61 on the Intelligence Index, a five-point gain over Grok 4.5 barely a month after that release and 23 points above Grok 4.3, putting it level with GPT-5.6 Sol and behind only Claude Opus 5 at 63 and Claude Fable 5 at 62. Pricing stays at $2 per million input tokens and $6 per million output tokens, which is more than sixty percent below Claude Opus 5 at $5 and $25 and GPT-5.6 Sol at $5 and $30. Holding headline pricing flat across a generation is unusual at the frontier, where capability gains have normally arrived with a price increase. The one quiet regression is cache-hit pricing, which rose from $0.30 to $0.50 per million tokens. Context stays at 500,000 tokens.
The model's strength is agentic rather than static reasoning. It posts a GDPval-AA v2 Elo of 1753, behind only Claude Opus 5 and statistically indistinguishable from Claude Fable 5 and Qwen3.8 Max given overlapping confidence intervals. It reaches 50.7 percent on the banking variant of the tau-cubed customer-service benchmark, essentially tied with Qwen3.8 Max at 51.3 percent, and 88.4 percent on Terminal-Bench version 2.1. Artificial Analysis measures cost per task at $0.84, the same as Kimi K3 at slightly higher intelligence, which places it on the intelligence-versus-cost Pareto frontier for every agentic evaluation in the index. The turn-efficiency numbers are the most striking part: roughly 53 turns and half a billion input tokens on average, against roughly 103 turns and two billion input tokens for Claude Opus 5 at maximum effort. A model that reaches a comparable answer in half the turns and a quarter of the input tokens carries a cost advantage well beyond its per-token sticker.
xAI attributes the gain to a longer supplemental pre-training run with curated model-generated reasoning data, an improved optimizer and recipe, regenerated supervised fine-tuning trajectories produced by Grok 4.5 across reasoning efforts and agent harnesses, and agentic reinforcement learning over kernel optimization, web development and computer-aided design environments. Practitioners on X report a confirmed 1.5-trillion-parameter model. The caveats are visible in xAI's own table: Grok 4.6 trails on DeepSWE version 1.1 at 65.9 percent against 73 percent for GPT-5.6 Sol and 70 percent for Fable 5, and badly on Terminal-Bench version 3.0 at 26 percent against 34.6 and 34.1 percent. It is available today in Cursor, Grok Build, the API, OpenRouter, Vercel and Cloudflare, with a fast variant at double the price. Elon Musk has already said Grok 4.7 has finished initial training.
- Artificial Analysis frames the release as cost efficiency first, calling out the flat $2/$6 pricing as the unusual move rather than the five-point index gain.
- xAI's own eval table concedes DeepSWE and Terminal-Bench v3.0 deficits that the Artificial Analysis writeup does not surface.
- Latent Space notes Cognition and Musk both position it as roughly the second-best knowledge-work model, and cites a confirmed 1.5T parameter count absent from the official post.
- Hacker News discussion (504 points, 458 comments) centered on the SpaceXAI rebrand following SpaceX's acquisition of xAI in February.