Humanity's Last Exam
2 published articles
Benchmarks & TestsFeatured2 min read
Benchmark Analysis
Humanity's Last Exam Opened With a 2.7% Score. The Best Models Still Can't Break 65%
Humanity's Last Exam was built by CAIS and Scale AI to succeed MMLU as the hardest general LLM benchmark. Two years on, even Claude Opus 5's leading 64.7% score sits well below the 90% human expert baseline.
2026-07-31
LLMs & Models4 min read
AI Research
Four minds, one answer: why AI that thinks differently beat the biggest models at humanity's hardest test
PoTRE breaks inference into four agents working in parallel, then reconciles their answers dynamically. It hit 49.92% on Humanity's Last Exam, beating the previous official best. The approach works with fewer tokens than scaled-up baselines.
2026-07-24