AI benchmark
5 published articles
Biosafety Benchmarks
The 50 percent ceiling on AI pathogen surveillance that should keep us up at night
BioSecBench-Surveillance tested sixteen AI agent configurations on 100 genomic surveillance tasks. The top performers managed only about 50 percent accuracy. The mistakes came not from picking the wrong workflow but from the surrounding choices - references, thresholds, filters - that humans take for granted.
2026-07-25
Agentic coding
Grok 4.5 just broke the coding agent leaderboard: the lead is real, the margins are tiny
Grok 4.5 now leads the SWE Marathon leaderboard, beating Claude 4 Opus and GPT-5. The benchmark tests real software engineering skills: bug fixes, feature additions, and code understanding across real repositories. The margin is slim, but the trend lines point toward a shrinking gap between what agents can do and what they need to do.
2026-07-20
Research analysis
AI document corruption in delegated workflows: what a new stress test reveals
The DELEGATE-52 benchmark evaluates AI systems on long-horizon delegated document editing tasks, finding that frontier models accumulate semantic fidelity loss of 19–34% over 20 iterations. Python workflows showed less than 1% degradation on average, but the study underscores that reliable long-horizon delegation remains an open challenge.
2026-07-04
Benchmark Deep Dive
ARC-AGI-2: The Benchmark That Measures Fluid Intelligence in AI Systems
ARC-AGI-2 tests AI systems on fluid intelligence through visual grid puzzles that can't be solved by memorization. Top frontier models now score 75-85%, but the grand prize of $700,000 remains unclaimed. Here's a deep dive into the benchmark's design, scoring, and current leaderboard.
2026-07-01
Open source
The first open-source web agent that doesn't need HTML just beat GPT-4o
MolmoWeb agents, available in 4B and 8B sizes, were trained on MolmoWebMix, a new dataset combining over 130,000 synthetic and human demonstrations. They surpass closed models and set new standards for open web agent research.
2026-04-10