on-policy distillation
3 published articles
LLM distillation
LLM distillation stalls when context goes static. Flux-OPD makes it evolve
Contexts can carry task preferences into LLM training, but once distilled into a student they add little supervision. Peking University's Flux-OPD keeps the context moving with the student and uses a conflict term to weight teacher corrections. The abstract reports gains over existing OPD paradigms without naming numbers.
2026-07-31
Agentic RL research
RL's sparse-reward blind spot meets SEED: agents write their own lessons
SEED (Self-Evolving On-Policy Distillation) lets an LLM analyze its own past trajectories, extract reusable natural-language skills from them in hindsight, and distill those lessons back into its policy during RL. The result: denser, on-policy supervision that improves success rates by up to 22% on long-horizon tasks.
2026-07-25
AI Research
Qwen just taught diffusion models a trick from the LLM playbook: RL beats supervised fine-tuning
The Qwen-Image-2.0-RL report details a post-training pipeline that combines RLHF, on-policy distillation, and composite reward models to improve text-to-image and image editing quality. The approach yields measurable gains across aesthetic quality, instruction following, and face identity preservation, borrowing techniques from LLM alignment research.
2026-07-20