LLMs & ModelsFeatured4 min read
Research
RL's oldest blind spot just met a fix that lets agents learn from their own mistakes
SEED (Self-Evolving On-Policy Distillation) lets an LLM analyze its own past trajectories, extract reusable natural-language skills from them in hindsight, and distill those lessons back into its policy during RL. The result: denser, on-policy supervision that improves success rates by up to 22% on long-horizon tasks.
2026-07-25