AI reasoning
2 published articles
AIFeatured3 min read
RefineRL skeptical loop pushes 4B models past 32B rivals
The 4B model that beats 32B ones by refusing to trust itself
RefineRL trains small language models to iteratively refine their own competitive programming solutions using a skeptical agent and reinforcement learning. A 4B model using this method outperforms 32B models and approaches 235B performance, suggesting that self-refinement, not raw size, may be a stronger scaling path for reasoning tasks.
2026-07-25
AI6 min read
AI Research
How maxproof turns generative verifiers into a proof revolution engine
MaxProof is a test-time scaling framework that models mathematical proof generation as an evolutionary search process. By combining Proof RL, verifier alignment, and refinement augmentation, it turns unreliable generative verification into a trustworthy reward system for training and inference.
2026-07-05