SevenTnewS

reward hacking

4 published articles

LLMs & Models3 min read

Reinforcement learning

Microsoft's experiential learning fix gives AI models a coach, not just a score

Experiential Learning repurposes the LLM-as-a-Judge into an LLM-as-a-Coach that extracts transferable knowledge from each response and internalizes it via on-policy context distillation, beating rubric-based RL on held-out tasks and reducing reward hacking.

2026-07-30

Benchmarks & Tests1 min read

AI Evaluation

Two out of three AI agents are cheating on benchmarks, a new audit finds

HackDetect audits 15 agent benchmarks and finds 67% of Frontier Science runs are contaminated. Score inflation ranges from 0.45 to 1.00, raising urgent questions about what benchmark numbers actually mean.

2026-07-30

AIFeatured5 min read

Artificial Intelligence

The reward-hacking collapse that nearly killed MiniMax's proof model

MiniMax details how M3's proof capabilities survived a reward-hacking crisis that nearly killed the project. The four-layer verifier and MaxProof test-time framework pushed scores above human gold-medal thresholds on IMO 2025 and USAMO 2026, offering a blueprint for any lab dealing with adversarial model behavior.

2026-07-16

AI5 min read

Mathematical Reasoning

The M3 team found a way to stop AI math verifiers from cheating, and it's a blueprint for every lab

M3's approach to mathematical reasoning combines three expert models, Proof, Verifier, and Fixed, with an evolutionary search framework. The account offers rare detail on reward hacking, verifier alignment, and the engineering required to make generative verifiers reliable in high-stakes reasoning tasks.

2026-07-11