SevenTnewS

Benchmark Analysis

MMLU Ruled AI Benchmarking for Years. Then Models Started Acing the Answer Key

MMLU defined a generation of AI benchmarking, but training-data contamination and a flawed answer key have pushed frontier labs toward dynamic successors like HLE and GPQA Diamond.

Emmanuel Fabrice Omgbwa Yasse

2026-07-28 · 2 min read

MMLU Ruled AI Benchmarking for Years. Then Models Started Acing the Answer Key
Sources : Évaluation et B…·MMLU — official…·Measuring Massi…

MMLU, the Massive Multitask Language Understanding benchmark, spent the early 2020s as the default yardstick for language model quality. Fifty-seven subjects, multiple-choice format, everything from abstract algebra to US foreign policy. It was easy to run, easy to compare, and easy to put in a press release.

That convenience is exactly what broke it. Once a benchmark's questions are public text, they eventually end up in someone's training corpus. Independent audits of multilingual test sets have found contamination rates as high as 91.8% in large models, and the pattern is consistent: the bigger the model, the more of the test it appears to have memorized rather than reasoned through. A model that has seen a question and its answer during pretraining does not need to solve anything. It just needs to recall.

Contamination is only half the problem. The test itself has holes. Audits of MMLU's math subset found roughly 2% of questions were simply wrong, mislabeled answers, ambiguous phrasing, or math that doesn't check out. Two percent sounds small until you realize every model's score is being compared against a ceiling that includes broken questions no one can actually answer correctly.

What saturation actually hides

By the mid-2020s, frontier models were clearing 88% to 90% on MMLU. That looks like triumph. It's closer to a measurement floor. When most of the field bunches up near the top of a test, the test stops telling you anything about who is actually better. A two-point gap between two models on a saturated benchmark could reflect real capability, or it could reflect which model got luckier with the flawed 2%.

This is the core argument researchers now make for retiring MMLU from serious evaluation: it can no longer distinguish genuine reasoning from statistical recall. A model can score well by having absorbed the internet, including the parts of the internet that happen to be MMLU.

What replaced it

The response wasn't to patch MMLU. It was to build tests that can't be memorized in the same way. Humanity's Last Exam was designed explicitly as an MMLU successor, with questions vetted to resist both web search and pretraining recall. GPQA Diamond does something similar for graduate-level science. Dynamic benchmarks that regenerate their question pools on a rolling basis, rather than shipping a fixed set once and reusing it for years, are becoming the norm precisely because a static answer key has a shelf life.

None of this means MMLU is worthless. It's still a reasonable sanity check, a model that scores badly on MMLU probably has real problems. But as a leaderboard metric, a number a lab uses to claim state of the art, its usefulness ran out around the same time everyone stopped failing it.

Get the tech essentials in 3 minutes every morning

One email, every weekday, with what actually matters in AI and tech.