AI benchmarks
19 published articles
Analysis
Benchmarks, not bargains: China's AI labs reach parity
Chinese AI labs Qwen, DeepSeek, MiniMax, and Kimi have matched or beaten US frontier models in reasoning, coding, and document parsing benchmarks. With aggressive pricing and open-weight strategies, they are no longer just cheap alternatives, they are real competitors.
2026-07-27
Open Source AI
Kimi K3 is the biggest open model ever. It's still not the best.
Kimi K3 is the largest open model at 2.8T parameters, but it fails to beat the best proprietary models on overall benchmarks. While it excels on specific coding and agentic tasks, the gap to frontier leaders like Claude Fable 5 and GPT 5.6 Sol exposes the limits of scaling without architectural and data breakthroughs.
2026-07-27
Artificial Intelligence
The AI that taught itself to revise: four agents, one puzzle, no consensus
Researchers propose ARCANA, a multi-agent system that cracks ARC-AGI-2 puzzles by splitting reasoning into four steps and closing the loop with reflective feedback. Each failure teaches the next attempt. The architecture improves reasoning efficiency, but questions about generalization and compute cost remain.
2026-07-19
Orchestrator swap
Perplexity swapped its orchestrator model and the cost-performance chart is brutal
Perplexity swapped Grok 4.5 into Computer as its orchestrator model, claiming top accuracy on its internal benchmark at roughly half the cost of Claude Opus 4.8. The move signals a broader shift: which model you pair with an agent matters as much as the model itself.
2026-07-14
Benchmark integrity
GPT-5.6 Sol almost cracked a physics benchmark built so AI couldn't cheat
GPT-5.6 Sol (max) leads CritPt, a new physics benchmark built from unpublished graduate-level research problems, scoring roughly 5 points above GPT-5.5 and 4 points ahead of Claude Fable 5, but even the winner solved only a third of the problems.
2026-07-13
Frontier AI Deployment
Anthropic split its most powerful model in two, and that changes how AI gets deployed
Anthropic's Claude Fable 5 brings Mythos-class capability to public users, while Mythos 5 remains trusted-access only. The deployment model, capability routing, fallback classifiers, and tiered access, represents a fundamental shift in how frontier AI is released and used.
2026-07-10
Artificial Intelligence
Ifbench reveals the instruction-following gap that other benchmarks miss
IFBench measures language models' ability to follow precise natural-language instructions. xAI's Grok models lead while Claude models lag, showing instruction following is a distinct capability from general intelligence.
2026-07-07