Labs & Research3 min read
Fisher-R1 | Hypothesis testing
AI agents run flawless statistics and still draw the wrong conclusions
Agents that automate hypothesis testing can execute analyses flawlessly yet reach wrong conclusions, because most benchmarks never check whether the p-value is valid. A 425-task benchmark quantifies the gap, and an RL-trained open-weight model closes a chunk of it.
2026-08-19