SevenTnewS

Benchmark Analysis

GPQA Diamond Was Built to Be Google-Proof. Frontier Models Are Now Clearing 95% Anyway

GPQA Diamond was designed so search-equipped humans can't reliably answer its expert-level science questions. Frontier models are now clearing 89% to 96%, pushing the benchmark toward the same saturation that retired MMLU.

Emmanuel Fabrice Omgbwa Yasse

2026-08-01 · Last updated: 2026-08-03 · 2 min read

GPQA Diamond Was Built to Be Google-Proof. Frontier Models Are Now Clearing 95% Anyway
Sources : Évaluation et B…·GPQA — official…·GPQA: A Graduat…

GPQA stands for Google-Proof Q&A, and the name is a design constraint, not marketing. Every question in the set was written by a PhD-level expert in physics, chemistry, or biology, then tested against non-expert humans who were given full, unrestricted access to search engines. If a search-equipped human couldn't reliably find the answer, the question qualified. The "Diamond" subset is the hardest tier: questions where even domain experts within the relevant field, working without time pressure, don't always agree.

That's a genuinely difficult bar to clear. It's also, increasingly, a bar frontier models are clearing anyway. On the current leaderboard, Claude Sonnet 5 scores 96.2%, GPT-5.6 Sol scores 94.6%, Gemini 3.1 Pro scores 94.3%, and Claude Opus 5 and Claude Mythos 5 both sit above 95%. Even the models further down the pack, Kimi K3 at 91.5%, GLM 5.2 at 90.8%, DeepSeek V4 Flash at 89.2%, are operating in a band that would have been unthinkable when GPQA launched.

A benchmark eating its own moat

This is the same saturation pattern that killed MMLU, just arriving faster because GPQA Diamond is a smaller, more concentrated test. When nearly every frontier model clears 89% or higher, the benchmark stops discriminating between them in any meaningful way. A 96.2% and a 89.2% look like a big gap on paper, but both scores describe models that are, in practical terms, correctly answering almost every question a domain PhD would also get right. The remaining errors cluster on the genuinely contested edge cases the Diamond tier was built around, which is a much narrower and noisier signal than the benchmark's original design intended.

Compare this to Humanity's Last Exam, where the same top-tier models top out around 65%. The two benchmarks are pulling in different directions: GPQA Diamond is running out of room at the top, while HLE still has 35 points of daylight between the best model and human expert performance. That contrast is itself useful information. It tells you GPQA Diamond has effectively become a confirmation test, a way to check a model clears a high floor of graduate-level science reasoning, rather than a discriminator between the current frontier and whatever comes next.

Why it's still cited constantly

Despite the saturation, GPQA Diamond hasn't been retired the way MMLU has. Part of that is inertia: it's cheap to run, well understood, and its expert-validated questions are harder to argue with than a multiple-choice trivia set. Part of it is that a model failing GPQA Diamond is still a real red flag, even if passing it no longer separates the field. Researchers increasingly treat it as a floor to clear rather than a ceiling to chase, with HLE and dynamic, contamination-resistant benchmarks doing the actual work of ranking frontier systems against each other.

Get the tech essentials in 3 minutes every morning

One email, every weekday, with what actually matters in AI and tech.