benchmark saturation
4 published articles
Special Report: AI Evaluation
The State of AI Benchmarking in 2026: Inside the Collapse of Static Tests and the Systems Built to Replace Them
A comprehensive tour of the 2026 AI benchmarking landscape: why MMLU, GSM8K, and HumanEval broke, how dynamic benchmarks and expert exams like HLE and GPQA Diamond replaced them, what agentic and jagged-intelligence testing reveals, and how human preference, LLM judges, and production observability now round out the full evaluation stack.
2026-08-03
Benchmark Analysis
GPQA Diamond Was Built to Be Google-Proof. Frontier Models Are Now Clearing 95% Anyway
GPQA Diamond was designed so search-equipped humans can't reliably answer its expert-level science questions. Frontier models are now clearing 89% to 96%, pushing the benchmark toward the same saturation that retired MMLU.
2026-08-01
Grade-School Math Benchmark Quality Audit
GSM8K Was Supposed to Test Grade-School Math. Up to 42% of Its Questions Don't
GSM8K became a standard test of arithmetic reasoning in LLMs, but audits have found error rates in its question set as high as 42%, undermining what its scores actually measure.
2026-07-29
Benchmark Analysis
MMLU Ruled AI Benchmarking for Years. Then Models Started Acing the Answer Key
MMLU defined a generation of AI benchmarking, but training-data contamination and a flawed answer key have pushed frontier labs toward dynamic successors like HLE and GPQA Diamond.
2026-07-28