data contamination
3 published articles
Special Report: AI Evaluation
The State of AI Benchmarking in 2026: Inside the Collapse of Static Tests and the Systems Built to Replace Them
A comprehensive tour of the 2026 AI benchmarking landscape: why MMLU, GSM8K, and HumanEval broke, how dynamic benchmarks and expert exams like HLE and GPQA Diamond replaced them, what agentic and jagged-intelligence testing reveals, and how human preference, LLM judges, and production observability now round out the full evaluation stack.
2026-08-03
Benchmark Analysis
LiveBench Refuses to Sit Still. That's the Whole Point
LiveBench replaces a fixed answer key with a continuously refreshed pool of coding problems, repositories, and prediction questions, sidestepping the contamination that undermined MMLU and GSM8K. It's part of a broader shift toward dynamic evaluation alongside LiveCodeBench, ForecastBench, and LLMEval-Fair.
2026-08-03
Benchmark Analysis
MMLU Ruled AI Benchmarking for Years. Then Models Started Acing the Answer Key
MMLU defined a generation of AI benchmarking, but training-data contamination and a flawed answer key have pushed frontier labs toward dynamic successors like HLE and GPQA Diamond.
2026-07-28