LLM agents
15 published articles
Long-horizon reasoning
The one framework that lets AI agents remember what they did 47 steps ago
PRO-LONG is a minimal framework that helps LLM agents retain and retrieve information across long sequences of actions. On the ARC-AGI-3 benchmark, it improved average pass rates by 18 percentage points across frontier models while using 4.2 to 5.8 times fewer tokens than existing harnesses. With Fable 5, it hit 97.4% best@2 at a total inference cost of $1,750.
2026-07-24
AI Research
The one training trick that stops AI agents from freezing in production
Current LLM agents crumble under real-world randomness. NoisyAgent exposes them to controlled noise during training, improving both robustness and general benchmark performance. The paper suggests the field has been overfitting to pristine conditions.
2026-07-17
AI research
Your AI agent passed by accident. SkillCoach grades the process, not the answer.
SkillCoach is a self-evolving rubric framework that evaluates and improves agentic skill-use by analyzing skill selection, following, composition, and reflection processes, providing better supervision than outcome-only metrics.
2026-07-06