SevenTnewS

benchmark saturation

4 published articles

Benchmarks & TestsFeatured8 min read

Special Report: AI Evaluation

The State of AI Benchmarking in 2026: Inside the Collapse of Static Tests and the Systems Built to Replace Them

A comprehensive tour of the 2026 AI benchmarking landscape: why MMLU, GSM8K, and HumanEval broke, how dynamic benchmarks and expert exams like HLE and GPQA Diamond replaced them, what agentic and jagged-intelligence testing reveals, and how human preference, LLM judges, and production observability now round out the full evaluation stack.

2026-08-03

Benchmarks & TestsFeatured2 min read

Benchmark Analysis

GPQA Diamond Was Built to Be Google-Proof. Frontier Models Are Now Clearing 95% Anyway

GPQA Diamond was designed so search-equipped humans can't reliably answer its expert-level science questions. Frontier models are now clearing 89% to 96%, pushing the benchmark toward the same saturation that retired MMLU.

2026-08-01

Benchmarks & TestsFeatured2 min read

Grade-School Math Benchmark Quality Audit

GSM8K Was Supposed to Test Grade-School Math. Up to 42% of Its Questions Don't

GSM8K became a standard test of arithmetic reasoning in LLMs, but audits have found error rates in its question set as high as 42%, undermining what its scores actually measure.

2026-07-29

Benchmarks & TestsFeatured2 min read

Benchmark Analysis

MMLU Ruled AI Benchmarking for Years. Then Models Started Acing the Answer Key

MMLU defined a generation of AI benchmarking, but training-data contamination and a flawed answer key have pushed frontier labs toward dynamic successors like HLE and GPQA Diamond.

2026-07-28