Benchmarks & Tests
Comparisons, performance tests and model evaluations.
23 published articles
Financial reasoning benchmark
LLMs know accounting formulas. FinIndices shows they can't apply them
LLMs can recite accounting formulas, but FinIndices, a benchmark built on real Chinese financial statements, shows they struggle to apply them. Gemini-3.1-Pro's accuracy fell from 70.70% to 38.22% when formula hints were removed, and generating tables actively degraded models' reasoning.
2026-08-20
AI in Education
AI tutors over-help by design. TutorMoments quantifies it
A benchmark built on 462 real tutoring sessions finds AI tutors default to doing the work for students. Explicit prompting helps, but models still struggle with the core judgment call: when to push, when to scaffold, when to step back.
2026-08-10
Voice Agents
Why Grok 4.1 Fast beats smarter models on a phone call
Overall leaderboards pick the wrong LLMs for voice calls. BenchLM's 2026 ranking puts latency first: Grok 4.1 Fast answers in 0.54s while Claude Opus 4.6, the top scorer, needs 1.78s. Fast models take the conversation; reasoning models stay on background tool calls.
2026-08-05
AI Evaluation
The old AI benchmarks broke. Here's what replaced them.
Static benchmarks like MMLU and GSM8K are saturated and contaminated. The industry has moved to dynamic regeneration, expert-level exams, and agentic task environments to get a real measure of AI capability.
2026-08-04
Special Report: AI Evaluation
The State of AI Benchmarking in 2026: Inside the Collapse of Static Tests and the Systems Built to Replace Them
A comprehensive tour of the 2026 AI benchmarking landscape: why MMLU, GSM8K, and HumanEval broke, how dynamic benchmarks and expert exams like HLE and GPQA Diamond replaced them, what agentic and jagged-intelligence testing reveals, and how human preference, LLM judges, and production observability now round out the full evaluation stack.
2026-08-03
Benchmarking
Frontier AI vision models fail at basic perception, new benchmark shows
PerceptionBench tests ten atomic visual capabilities across 3,000 questions. No frontier model cracked 60 percent, and similar overall scores mask wildly different weakness profiles.
2026-08-03
Benchmark Analysis
LiveBench Refuses to Sit Still. That's the Whole Point
LiveBench replaces a fixed answer key with a continuously refreshed pool of coding problems, repositories, and prediction questions, sidestepping the contamination that undermined MMLU and GSM8K. It's part of a broader shift toward dynamic evaluation alongside LiveCodeBench, ForecastBench, and LLMEval-Fair.
2026-08-03
Benchmark Analysis
5 Million Votes, 25 Elo Points: Inside the Chatbot Arena Leaderboard Nobody Can Shake
LMSYS Chatbot Arena ranks models by blind human preference on an Elo scale, with nearly five million votes now packing the top labs into a 25-point band. The LLM-as-a-Judge methods that scale this kind of evaluation carry their own documented biases toward verbosity and position.
2026-08-02
Benchmark Analysis
On SWE-bench Verified, Top Models Hit 96%. On Private Enterprise Code, They Barely Clear 23%
SWE-bench Verified makes frontier models look close to solving real-world software engineering, with top scores above 95%. SWE-bench Pro, run on private enterprise repositories, drops those same models to 23% or lower, exposing how much of the Verified score depended on public data exposure.
2026-08-02
Benchmark Analysis
GPQA Diamond Was Built to Be Google-Proof. Frontier Models Are Now Clearing 95% Anyway
GPQA Diamond was designed so search-equipped humans can't reliably answer its expert-level science questions. Frontier models are now clearing 89% to 96%, pushing the benchmark toward the same saturation that retired MMLU.
2026-08-01
Benchmarks
AI desktop agents fail before-after test 35% of the time
DDB tests ordering and before-after pair tasks across 2,013 instances. The top model hit 65.1% exact match on non-decoy sequences and 65.7% with decoys, exposing a gap in how agents verify state changes.
2026-08-01
Benchmark Analysis
Humanity's Last Exam Opened With a 2.7% Score. The Best Models Still Can't Break 65%
Humanity's Last Exam was built by CAIS and Scale AI to succeed MMLU as the hardest general LLM benchmark. Two years on, even Claude Opus 5's leading 64.7% score sits well below the 90% human expert baseline.
2026-07-31