dynamic benchmarks
2 published articles
Benchmarks & TestsFeatured5 min read
AI Evaluation
The old AI benchmarks broke. Here's what replaced them.
Static benchmarks like MMLU and GSM8K are saturated and contaminated. The industry has moved to dynamic regeneration, expert-level exams, and agentic task environments to get a real measure of AI capability.
2026-08-04
Benchmarks & TestsFeatured2 min read
Benchmark Analysis
LiveBench Refuses to Sit Still. That's the Whole Point
LiveBench replaces a fixed answer key with a continuously refreshed pool of coding problems, repositories, and prediction questions, sidestepping the contamination that undermined MMLU and GSM8K. It's part of a broader shift toward dynamic evaluation alongside LiveCodeBench, ForecastBench, and LLMEval-Fair.
2026-08-03