LLMs & Models5 min read
AI Research
Distributional RL heads make risk claims that are mostly false. A new audit proves it.
A large-scale audit shows that distributional RL agents systematically fabricate risk trade-offs at the very states where practitioners would most trust them. Across QR-DQN, C51, and IQN on MinAtar, zero of the strongest claims were confirmable, and acting on the agents' CVaR advice sometimes performed significantly worse than chance.
2026-07-29