arXiv cs.LGAugust 17, 2026
Designing Reinforcement Learning for Diffusion Models: A Unified Path-Space View
Excerpt
arXiv:2608.14430v1 Announce Type: new Abstract: Reinforcement learning (RL) post-training provides a direct way to align diffusion models with human preferences and task-specific rewards. However, current RL algorithms for diffusion models remain fragmented: reverse-trajectory methods rely on discretized likelihood ratios, whereas forward-matching methods train on reward-labeled noising versions of the rollout samples. This paper shows that these seemingly different losses arise from a single pa