SevenTnewS

Benchmark Analysis

GSM8K Was Supposed to Test Grade-School Math. Up to 42% of Its Questions Don't

GSM8K became a standard test of arithmetic reasoning in LLMs, but audits have found error rates in its question set as high as 42%, undermining what its scores actually measure.

Emmanuel Fabrice Omgbwa Yasse

2026-07-29 · 2 min read

GSM8K Was Supposed to Test Grade-School Math. Up to 42% of Its Questions Don't
Sources : Évaluation et B…·GSM8K — officia…·Training Verifi…

GSM8K, short for Grade School Math 8K, is a set of roughly 8,000 word problems pitched at a middle-school level: trains leaving stations, people splitting bills, that kind of thing. For a few years it was the standard way to check whether a language model could chain together several arithmetic steps without losing the thread. Solving GSM8K reliably was treated as evidence of real multi-step reasoning, as opposed to pattern matching.

Then independent audits started reading the questions closely instead of just running models against them. What they found was a benchmark with a quality problem, not a difficulty problem. Error rates in GSM8K's item pool have been estimated as high as 42%: questions with ambiguous wording, missing information, or reference answers that don't match any defensible reading of the problem. A model marked "wrong" on one of these isn't necessarily reasoning worse than a model marked "right." It might just be reading the question more carefully.

Why this matters more than it sounds like it should

A 2% error rate is noise. A 42% error rate is a different benchmark than the one everyone thinks they're running. If nearly half the test items are unreliable, then the score itself stops being a clean measurement of arithmetic reasoning and starts reflecting something closer to how well a model guesses what a flawed rubric wants to hear.

This compounds with the same contamination problem that hit MMLU. GSM8K's questions are old enough, and public enough, that large models have almost certainly seen many of them, or close paraphrases, during pretraining. Combine memorization with a shaky answer key, and a high GSM8K score tells you less than it used to about whether a model can actually do the arithmetic, and more about whether it happened to land on the same interpretation of a badly worded problem that the dataset's authors did.

What labs use instead

Math evaluation didn't disappear, it moved to harder and more carefully audited territory. Competition-style math benchmarks pulled from recent olympiad and contest problems are harder to contaminate simply because the problems are newer than most training cutoffs. Multi-step reasoning is increasingly tested inside broader graduate-level exams like GPQA Diamond and Humanity's Last Exam, where math is one discipline among many rather than the entire test, and where the item-quality bar is explicitly higher: HLE's own verification process was built specifically to catch the kind of errors that plagued GSM8K.

None of this means GSM8K is useless for a quick sanity check. A model that fails basic word problems clearly has a problem. But treating a GSM8K percentage as a precise measure of reasoning ability, the way papers did for years, assumes an answer key that turns out to be a lot less solid than it looked.

Get the tech essentials in 3 minutes every morning

One email, every weekday, with what actually matters in AI and tech.