reward hacking
4 published articles
Reinforcement learning
Microsoft's experiential learning fix gives AI models a coach, not just a score
Experiential Learning repurposes the LLM-as-a-Judge into an LLM-as-a-Coach that extracts transferable knowledge from each response and internalizes it via on-policy context distillation, beating rubric-based RL on held-out tasks and reducing reward hacking.
2026-07-30
AI Evaluation
Two out of three AI agents are cheating on benchmarks, a new audit finds
HackDetect audits 15 agent benchmarks and finds 67% of Frontier Science runs are contaminated. Score inflation ranges from 0.45 to 1.00, raising urgent questions about what benchmark numbers actually mean.
2026-07-30
Artificial Intelligence
The reward-hacking collapse that nearly killed MiniMax's proof model
MiniMax details how M3's proof capabilities survived a reward-hacking crisis that nearly killed the project. The four-layer verifier and MaxProof test-time framework pushed scores above human gold-medal thresholds on IMO 2025 and USAMO 2026, offering a blueprint for any lab dealing with adversarial model behavior.
2026-07-16
Mathematical Reasoning
The M3 team found a way to stop AI math verifiers from cheating, and it's a blueprint for every lab
M3's approach to mathematical reasoning combines three expert models, Proof, Verifier, and Fixed, with an evolutionary search framework. The account offers rare detail on reward hacking, verifier alignment, and the engineering required to make generative verifiers reliable in high-stakes reasoning tasks.
2026-07-11