arXiv cs.LGOctober 2, 2026
iADD: Improving Alignment and Diversity in Diffusion Policy Optimization
Excerpt
arXiv:2610.01789v1 Announce Type: new Abstract: Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we provide better tradeoffs through careful theoretical considerations and method design. We analyze the theoretical framework and mathematically demonstrate