SevenTnewS

Grok 4.5

4 published articles

xAI / Grok3 min read

Ecosystem play

Grok 4.5's biggest news might be how it was built

xAI's Grok 4.5 is now on grok.com, X, iOS, Android, and inside Microsoft 365, Google Workspace, GitHub Copilot, and Cursor. The model, built with agent-constructed training data, leads the SWE Marathon benchmark and introduces a voice upgrade.

2026-08-01

Benchmarks & TestsFeatured3 min read

Agentic coding

Grok 4.5 just broke the coding agent leaderboard: the lead is real, the margins are tiny

Grok 4.5 now leads the SWE Marathon leaderboard, beating Claude 4 Opus and GPT-5. The benchmark tests real software engineering skills: bug fixes, feature additions, and code understanding across real repositories. The margin is slim, but the trend lines point toward a shrinking gap between what agents can do and what they need to do.

2026-07-20

AI1 min read

Orchestrator swap

Perplexity swapped its orchestrator model and the cost-performance chart is brutal

Perplexity swapped Grok 4.5 into Computer as its orchestrator model, claiming top accuracy on its internal benchmark at roughly half the cost of Claude Opus 4.8. The move signals a broader shift: which model you pair with an agent matters as much as the model itself.

2026-07-14

xAI / GrokFeatured3 min read

Grok 4.5

Cursor's Grok 4.5 was built by AI agents, not humans. That's the real story.

Cursor's Grok 4.5 is a Mixture-of-Experts model built using reinforcement learning in environments created by earlier AI agents, not humans. It handles complex, long-duration tasks across software engineering, data science, finance, and law, and it's available now.

2026-07-10