Grade-School Math Benchmark Quality Audit
GSM8K Was Supposed to Test Grade-School Math. Up to 42% of Its Questions Don't
GSM8K became a standard test of arithmetic reasoning in LLMs, but audits have found error rates in its question set as high as 42%, undermining what its scores actually measure.

GSM8K stands for Grade School Math 8K. It's about 8,000 word problems at a middle-school level: trains leaving stations, splitting dinner bills, that kind of thing. For a few years, if you wanted to test whether a language model could do real multi-step arithmetic, this was the benchmark. Researchers treated a strong GSM8K score as proof the model was doing actual reasoning, not just pattern matching.
Then independent researchers started reading the questions instead of just running models against them, echoing a trend seen in other benchmark integrity investigations. What they found was a quality problem, not a difficulty problem. Some estimates put the error rate as high as 42%: questions with ambiguous wording, missing information, or reference answers that don't line up with any reasonable reading. If a model is marked wrong on one of these, it may not be reasoning worse than a model that got it right. It might just be reading the question more carefully.
Why this matters more than it sounds like it should
A 2% error rate is acceptable noise, but 42% is a different benchmark than the one everyone thought they were running. When nearly half the items are unreliable, the score stops being a clean measurement of arithmetic reasoning and starts reflecting something closer to how well a model guesses what a flawed rubric wants.
This compounds with the same contamination problem that hit MMLU. GSM8K's questions are old enough, and public enough, that large models have almost certainly seen many of them, or close paraphrases, during pretraining. Combine memorization with a shaky answer key, and a high GSM8K score tells you less than it used to about whether a model can actually do the arithmetic, and more about whether it happened to land on the same interpretation of a badly worded problem that the dataset's authors did.
What labs use instead
Math evaluation didn't disappear. It moved to harder, more carefully audited territory. Competition-style benchmarks drawn from recent olympiad problems are harder to contaminate because those problems are newer than most training cutoffs. Multi-step reasoning is increasingly tested inside broader graduate-level exams like GPQA Diamond and Humanity's Last Exam, where math is one discipline among many. The item-quality bar is higher there: HLE's verification process was designed specifically to catch the sort of errors that plague GSM8K. This careful design matters because benchmarks can fail to distinguish description from genuine reasoning.
None of this means GSM8K is useless for a quick sanity check. A model that fails basic word problems clearly has a problem. But treating a GSM8K percentage as a precise measure of reasoning ability, the way papers did for years, assumes an answer key that turns out to be a lot less solid than it looked.
Get the tech essentials in 3 minutes every morning
One email, every weekday, with what actually matters in AI and tech.