Benchmark Analysis
MMLU Ruled AI Benchmarking for Years. Then Models Started Acing the Answer Key
MMLU defined a generation of AI benchmarking, but training-data contamination and a flawed answer key have pushed frontier labs toward dynamic successors like HLE and GPQA Diamond.

MMLU, the Massive Multitask Language Understanding benchmark, was the default yardstick for language model quality in the early 2020s. Fifty-seven subjects, multiple-choice format, everything from abstract algebra to US foreign policy. Running it was easy; comparing results across labs was trivial. That combination made it perfect for press releases.
That convenience is exactly what broke it. Once a benchmark's questions are public text, they eventually end up in someone's training corpus. Independent audits of multilingual test sets have found contamination rates as high as 91.8% in large models, and the pattern is consistent: the bigger the model, the more of the test it appears to have memorized rather than reasoned through. A model that has seen a question and its answer during pretraining does not need to solve anything. It just needs to recall.
Contamination is only half the problem. The test itself has holes. Audits of MMLU's math subset found roughly 2% of questions were simply wrong, mislabeled answers, ambiguous phrasing, or math that doesn't check out. Other benchmarks have fared worse: audits of GSM8K found error rates as high as 42% of its questions, per the GSM8K audit. Two percent sounds small until you realize every model's score is being compared against a ceiling that includes broken questions no one can actually answer correctly.
What saturation actually hides
By the mid-2020s, frontier models were clearing 88% to 90% on MMLU. That looks like triumph. It's closer to a measurement floor. When most of the field bunches up near the top of a test, the test stops telling you anything about who is actually better, as multiple case studies of benchmark blind spots have shown, see the analysis of benchmarks missing real-world performance. A two-point gap between two models on a saturated benchmark could reflect real capability, or it could reflect which model got luckier with the flawed 2%.
This is the core argument researchers now make for retiring MMLU from serious evaluation: it can no longer distinguish genuine reasoning from statistical recall. A model can score well by having absorbed the internet, including the parts of the internet that happen to be MMLU.
What replaced it
The response wasn't to patch MMLU. It was to build tests that can't be memorized in the same way. Humanity's Last Exam was designed explicitly as an MMLU successor, with questions vetted to resist both web search and pretraining recall. GPQA Diamond does something similar for graduate-level science. Dynamic benchmarks that regenerate their question pools on a rolling basis, rather than shipping a fixed set once and reusing it for years, are becoming the norm precisely because a static answer key has a shelf life, a lesson that extends to agent evaluation, as new agent benchmarks show.
None of this means MMLU is worthless. It's still a reasonable sanity check, a model that scores badly on MMLU probably has real problems. But as a leaderboard metric, a number a lab uses to claim state of the art, its usefulness ran out around the same time everyone stopped failing it.
Get the tech essentials in 3 minutes every morning
One email, every weekday, with what actually matters in AI and tech.