arXiv cs.LGAugust 17, 2026
Latent Reward Registers for Diffusion Preference Alignment
Excerpt
arXiv:2608.03929v3 Announce Type: replace Abstract: Aligning diffusion models with human preferences usually relies on a sparse terminal reward evaluated on the final generated samples, which creates a severe temporal credit-assignment problem across the denoising process. We propose Latent Reward Registers, a mechanism that estimates terminal preference directly from intermediate noisy latents. Learnable, position-free register tokens are appended as an auxiliary read path to a frozen Diffusion