arXiv preprint
4 published articles
AI Research
Muon groks modular addition faster, then its solutions collapse
Muon-trained transformers grok modular addition faster than AdamW, then lose the solution in every configuration tested. The paper pins the collapse to the representation-readout interface, where freezing either parameter group prevents it. Fourier analysis shows the task circuit survives and simply gets outvoted.
2026-08-19
Fisher-R1 | Hypothesis testing
AI agents run flawless statistics and still draw the wrong conclusions
Agents that automate hypothesis testing can execute analyses flawlessly yet reach wrong conclusions, because most benchmarks never check whether the p-value is valid. A 425-task benchmark quantifies the gap, and an RL-trained open-weight model closes a chunk of it.
2026-08-19
Clinical AI: heart-failure phenotyping preprint
nMAS automates heart-failure EHR features, tested only on 500 dummy patients
nMAS, a multi-agent pipeline, generated 132 structured and 70 rubric-scored features from 500 dummy patient records, lifting held-out HFrEF phenotyping AUROC from 0.895 to 0.963. The preprint has not been validated on real patient data.
2026-08-13
AI Research
A 4.5 point score jump on MATH-500 from a monitoring controller that catches wandering models
New research proposes an external monitoring controller for quantized small language models that detects repetitive or degenerating reasoning paths and triggers a rollback and constrained re-decoding. Accuracy improved by 4.5 percentage points on a broad evaluation set, but the authors stress the findings are not confirmatory.
2026-07-30