AI alignment
2 published articles
LLMs & Models3 min read
Reinforcement learning
Microsoft's experiential learning fix gives AI models a coach, not just a score
Experiential Learning repurposes the LLM-as-a-Judge into an LLM-as-a-Coach that extracts transferable knowledge from each response and internalizes it via on-policy context distillation, beating rubric-based RL on held-out tasks and reducing reward hacking.
2026-07-30
Labs & ResearchFeatured4 min read
AI Research
RDPO: The two-word fix for reinforcement learning's self-sabotage problem
Multi-task reinforcement learning has a dirty secret: the reward signals that drive alignment often work against each other. An international team just published a method that prevents the system from fighting itself, and it costs almost nothing to add to existing pipelines.
2026-07-19