Artificial Intelligence
Grok 4.6 ties GPT-5.6 Sol at 61, then the component scores split
xAI's Grok 4.6 ties GPT-5.6 Sol at 61 on the Artificial Analysis Intelligence Index, a composite of nine benchmarks. The breakdown is lopsided: wins on knowledge work and most coding tests, losses on DeepSWE and Terminal-Bench. Available now in Cursor and Grok Build, from $2 per million input tokens.
Emmanuel Fabrice Omgbwa Yasse AI-assisted
2026-08-13 · 4 min read

xAI's new release makes the case that frontier models should be judged on how long they can sustain a task, not how fast they answer a single prompt. Grok 4.6, out today, is built for work that runs for many steps: researching a topic, analyzing information, working across a codebase, carrying an idea through to a working application. xAI says the model stays with complex tasks across those trajectories and, on longer runs, starts checking its own work before moving on. The appetite for longer, less supervised runs is already a theme on the consumer side, where Grok now handles scheduled and email-triggered jobs without supervision.
The headline figure is a 61 on the Artificial Analysis Intelligence Index, a composite of nine benchmarks. That puts Grok 4.6 level with GPT-5.6 Sol and one point behind Fable 5 at 62. Measured against xAI's own prior release, the gain is clear: Grok 4.5 scored 56 on the same index.
The composite is where the parity lives. The component scores xAI published tell a more divided story. Grok 4.6 beats GPT-5.6 Sol on six of the listed benchmarks: GDPVal-AA v2 (1753 to 1728), CursorBench v3.2 (69.9% to 67.2%), FrontierCode v1.1 in its extended form (61.3% to 60.6%), APEX-Agents (57.5% to 56.7%), AA-Briefcase (1577 to 1502), and the Harvey LAB evaluation (Vals), where it posts 15.8% to its rival's 2.5%. On APEX-SWE, xAI reports 56.4% and GPT-5.6 Sol has no published score. The losses are concentrated in two rows: DeepSWE v1.1 at 65.9% against 73%, and Terminal-Bench v3.0 at 26% against 34.6%.
Those two deficits are the ones that matter for a release pitched at long-running agents. Both benchmarks reward sustained execution over a long session, which is exactly the territory the model is supposed to own. The trajectory is encouraging: DeepSWE climbs from 54% on Grok 4.5 to 65.9%, Terminal-Bench from 15.7% to 26%, and APEX-Agents from 47.1% to 57.5%, the three largest percentage-point gains on xAI's scorecard. But GPT-5.6 Sol still holds a seven-to-nine point edge on the two execution tests, and Fable 5 runs four to eight points ahead. The trait also explains why Alibaba's model coding alone for 16 days became a headline, and it makes these two rows a direct test case for the question Beacon raises about when tools are actually needed and when they hurt.
The scorecard xAI published
| Benchmark | Grok 4.6 | Grok 4.5 | GPT-5.6 Sol | Fable 5 |
|---|---|---|---|---|
| AA Intelligence Index | 61 | 56 | 61 | 62 |
| GDPVal-AA v2 | 1753 | 1526 | 1728 | 1741 |
| CursorBench v3.2 | 69.9% | 66.7% | 67.2% | 70.5% |
| DeepSWE v1.1 | 65.9% | 54% | 73% | 70% |
| FrontierCode v1.1 (Extended) | 61.3% | 56.6% | 60.6% | 63.6% |
| APEX-Agents | 57.5% | 47.1% | 56.7% | 59.2% |
| Terminal-Bench v3.0 | 26% | 15.7% | 34.6% | 34.1% |
| APEX-SWE | 56.4% | 53.6% | n/a | 58.8% |
| AA-Briefcase | 1577 | 1313 | 1502 | 1574 |
| Harvey LAB (Vals) | 15.8% | 12.9% | 2.5% | 11.3% |
Best score per row is in bold. Competitor figures are drawn from each developer's published system cards or benchmark leaderboards, per xAI's release.
How Grok 4.6 was trained: the self-built loop continues
The training story extends the pattern xAI set with Grok 4.5, whose training environments were built by autonomous agents rather than human engineers. Cursor, which trained 4.5 alongside xAI, disclosed that detail at the time, and it became the real story of that release. For 4.6, xAI describes a longer supplemental training run than 4.5's, fed with curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and an improved optimizer and recipe. Grok 4.5 then generated the SFT trajectories across reasoning efforts, agent harnesses, STEM, software engineering, and knowledge work, with model-based checks filtering out bad traces. The RL stage that followed covered agentic tasks including general coding and domain-specific environments for kernel optimization, web development, and computer-aided design.
The model-level results match the pitch. xAI says Grok 4.6 is especially strong at turning a broad product idea into a working first version, researching unfamiliar domains, structuring an application, implementing core interactions, and refining through several rounds of feedback. It also reports stronger first passes on visual and interactive work than Grok 4.5 typically produced, with a project's structure and visual language established in one pass.
Where it runs and what it costs
Grok 4.6 is live today in Grok Build and Cursor, through the xAI API, and via partners including OpenRouter, Vercel, and Cloudflare. The multi-partner rollout echoes Grok 4.5's launch across Microsoft 365, Google Workspace, and GitHub Copilot. Pricing starts at $2 per million input tokens and $6 per million output tokens, with a fast variant at twice the price. The rate sits far above the $0.07-per-million pricing that went viral for Zhipu's GLM-5.2. To get the model into developers' hands quickly, xAI is doubling included usage in Grok Build and Cursor for the first week.
On safety, xAI says safeguards were calibrated alongside the model's expanded capabilities, backed by its widest-ever suite of pre-deployment testing plus post-deployment and third-party testing. The release explicitly keeps legitimate work like vulnerability patching and accelerating engineering design cycles in scope. That stance carries weight given xAI's pivot toward defense partners.
The composite tie with GPT-5.6 Sol is real, and the scorecard underneath is narrower than the framing suggests. Grok 4.6 wins the work xAI aimed it at: knowledge-heavy research and interactive builds. It trails on the longest execution tasks. The gaps on DeepSWE and Terminal-Bench run seven to nine points behind GPT-5.6 Sol. It is a deficit small enough to close in a point release, but large enough to matter for exactly the workloads this model is being sold for.
- Source : xAI: Grok 4.6 release announcement
Get the tech essentials in 3 minutes every morning
One email, every weekday, with what actually matters in AI and tech.