MLE-Bench
2 published articles
AI Agents4 min read
AI research automation, arXiv, September 2026
DisCo reports a 134.3% MLE-bench gain with its backbone held fixed
DisCo distills 1,000 ML repositories into more than 5,000 reusable skills, and reports large benchmark gains with the model, harness and execution budget held fixed. The figures are self-reported and unreplicated.
2026-09-16
Labs & Research4 min read
Frontis-MA1 / OpenMLE
The AI that improves AI now runs on one RTX 4090
FrontisAI's OpenMLE stack and 35B Frontis-MA1 agent turn recursive self-improvement into a reproducible experiment. On a single RTX 4090, the model scores 71.21% on MLE-Bench Lite, ahead of GPT-5.5 + Codex and close to much larger frontier models.
2026-07-31