competitive programming
2 published articles
AIFeatured3 min read
RefineRL skeptical loop pushes 4B models past 32B rivals
The 4B model that beats 32B ones by refusing to trust itself
RefineRL trains small language models to iteratively refine their own competitive programming solutions using a skeptical agent and reinforcement learning. A 4B model using this method outperforms 32B models and approaches 235B performance, suggesting that self-refinement, not raw size, may be a stronger scaling path for reasoning tasks.
2026-07-25
LLMs & Models2 min read
Competitive programming
NousCoder-14B just opened the coding RL black box that OpenAI and DeepMind keep locked
Nous Research drops NousCoder-14B, a 14B competitive programming model with a fully open RL pipeline. The 68% Codeforces solve rate is notable, but the real story is that anyone can now replicate the stack.
2026-07-16