Benchmark Analysis
Humanity's Last Exam Opened With a 2.7% Score. The Best Models Still Can't Break 65%
Humanity's Last Exam was built by CAIS and Scale AI to succeed MMLU as the hardest general LLM benchmark. Two years on, even Claude Opus 5's leading 64.7% score sits well below the 90% human expert baseline.

When Humanity's Last Exam (HLE) launched in early 2025, GPT-4o scored 2.7%. Claude 3.5 Sonnet scored 4.1%. The human expert baseline sits around 90%. That gap wasn't an accident of a badly calibrated test, it was the point. HLE was built by the Center for AI Safety and Scale AI specifically because MMLU had stopped being hard enough to mean anything, and the project was formal enough to be published in Nature.
The exam has 2,500 questions, written by specialists across more than a hundred academic disciplines, and it's structured to resist the two things that killed the previous generation of benchmarks: web lookup and training-set memorization. There is no essay format to game with fluent-sounding filler. Answers are either short and exact (76% of questions) or multiple choice with distractors sophisticated enough to punish shallow pattern matching (24%). Fourteen percent of items require reading a diagram, a chart, or an image alongside the text, so a model can't get by on language alone.
Where the questions come from
The subject distribution tells you what the test actually values: mathematics makes up 41% of the question bank, biology and medicine 11%, computer science and AI 10%, physics 9%, chemistry 7%, humanities and social sciences 9%, engineering 4%, and assorted specialist domains the rest. It's a heavily quantitative exam by design, closer to a doctoral qualifying exam than a general-knowledge quiz.
The current leaderboard
As of the latest published scores, the field looks like this:
- Claude Opus 5: 64.7%
- Claude Mythos 5: 64.5%
- Claude Opus 4.8: 57.9%
- Claude Sonnet 5: 57.4%
- Kimi K3: 56.0%
- GLM 5.2: 54.7%
- Claude Fable 5: 53.3%
- Hy3 (Tencent): 53.2%
- DeepSeek V4 Flash: 51.6%
- GPT-5.6 Sol: 47.2%
- Gemini 3.1 Pro: 44.4%
- GPT-5.5 Pro: 43.1%
Two things stand out. First, the models with dedicated reasoning modes, ones that spend extra inference time planning and checking their own work before answering, dominate the top of the list. Raw parameter count stopped being the deciding factor once HLE became the reference test; what separates a 65% score from a 45% score increasingly looks like whether a model can catch its own mistakes mid-answer. Second, even the leader is still 25 points short of the human expert baseline, on a test explicitly engineered to be hard to game.
The exam wasn't perfect either
HLE's own early version had errors, an audit under the HLE-Verified initiative found that up to 30% of the chemistry and biology subsection contained ambiguous prompts or incorrect reference answers, forcing a component-by-component review. Rather than let that undermine the benchmark the way undiscovered errors undermined MMLU and GSM8K, the project built HLE-Verified to fix it and HLE-Rolling to keep adding vetted questions from the scientific community on an ongoing basis. That willingness to audit and republish, instead of shipping a fixed set once and defending it forever, is part of why HLE has held its position as the field's hardest general benchmark rather than following its predecessors into saturation.
Get the tech essentials in 3 minutes every morning
One email, every weekday, with what actually matters in AI and tech.