M3
6 published articles
Artificial Intelligence
The reward-hacking collapse that nearly killed MiniMax's proof model
MiniMax details how M3's proof capabilities survived a reward-hacking crisis that nearly killed the project. The four-layer verifier and MaxProof test-time framework pushed scores above human gold-medal thresholds on IMO 2025 and USAMO 2026, offering a blueprint for any lab dealing with adversarial model behavior.
2026-07-16
Mathematical Reasoning
The M3 team found a way to stop AI math verifiers from cheating, and it's a blueprint for every lab
M3's approach to mathematical reasoning combines three expert models, Proof, Verifier, and Fixed, with an evolutionary search framework. The account offers rare detail on reward hacking, verifier alignment, and the engineering required to make generative verifiers reliable in high-stakes reasoning tasks.
2026-07-11
AI Labs & Research
MiniMax's M3 just wrote its own CUDA kernel, and opened the code
MiniMax M3 scores 83.5 on BrowseComp, edges past Opus 4.7, and handles up to 1M tokens natively. In a remarkable autonomy test, it self-optimized a GPU kernel from 7.6% to 71.3% peak utilization without human intervention.
2026-07-09
Artificial Intelligence
MiniMax just shipped a model for every AI job you can name
Chinese AI startup MiniMax launches M3, Hailuo 2.3, MiniMax Code, and new speech/music models, broadening its product lineup in a competitive landscape.
2026-07-06
Artificial Intelligence
MiniMax's M3 just beat Opus 4.7 at browsing, trained itself, and never asked for help
MiniMax M3 delivers a 9.4x CUDA kernel speedup, beats Opus 4.7 on BrowseComp, and autonomously replicated an ICLR paper. All in an open-weight package, and it never asked for help.
2026-07-06
AI Research
How maxproof turns generative verifiers into a proof revolution engine
MaxProof is a test-time scaling framework that models mathematical proof generation as an evolutionary search process. By combining Proof RL, verifier alignment, and refinement augmentation, it turns unreliable generative verification into a trustworthy reward system for training and inference.
2026-07-05