LLM-as-a-judge
3 published articles
AI Evaluation
The old AI benchmarks broke. Here's what replaced them.
Static benchmarks like MMLU and GSM8K are saturated and contaminated. The industry has moved to dynamic regeneration, expert-level exams, and agentic task environments to get a real measure of AI capability.
2026-08-04
Special Report: AI Evaluation
The State of AI Benchmarking in 2026: Inside the Collapse of Static Tests and the Systems Built to Replace Them
A comprehensive tour of the 2026 AI benchmarking landscape: why MMLU, GSM8K, and HumanEval broke, how dynamic benchmarks and expert exams like HLE and GPQA Diamond replaced them, what agentic and jagged-intelligence testing reveals, and how human preference, LLM judges, and production observability now round out the full evaluation stack.
2026-08-03
Benchmark Analysis
5 Million Votes, 25 Elo Points: Inside the Chatbot Arena Leaderboard Nobody Can Shake
LMSYS Chatbot Arena ranks models by blind human preference on an Elo scale, with nearly five million votes now packing the top labs into a 25-point band. The LLM-as-a-Judge methods that scale this kind of evaluation carry their own documented biases toward verbosity and position.
2026-08-02