RLHF
2 published articles
LLMs & ModelsFeatured5 min read
AI Research
Qwen just taught diffusion models a trick from the LLM playbook: RL beats supervised fine-tuning
The Qwen-Image-2.0-RL report details a post-training pipeline that combines RLHF, on-policy distillation, and composite reward models to improve text-to-image and image editing quality. The approach yields measurable gains across aesthetic quality, instruction following, and face identity preservation, borrowing techniques from LLM alignment research.
2026-07-20
Labs & ResearchFeatured4 min read
AI Research
RDPO: The two-word fix for reinforcement learning's self-sabotage problem
Multi-task reinforcement learning has a dirty secret: the reward signals that drive alignment often work against each other. An international team just published a method that prevents the system from fighting itself, and it costs almost nothing to add to existing pipelines.
2026-07-19