Benchmarks & TestsFeatured3 min read
Agentic coding
Grok 4.5 just broke the coding agent leaderboard: the lead is real, the margins are tiny
Grok 4.5 now leads the SWE Marathon leaderboard, beating Claude 4 Opus and GPT-5. The benchmark tests real software engineering skills: bug fixes, feature additions, and code understanding across real repositories. The margin is slim, but the trend lines point toward a shrinking gap between what agents can do and what they need to do.
2026-07-20