← Back to all articles
arXiv cs.LGAugust 17, 2026

Latent Reward Registers for Diffusion Preference Alignment

Excerpt

arXiv:2608.03929v3 Announce Type: replace Abstract: Aligning diffusion models with human preferences usually relies on a sparse terminal reward evaluated on the final generated samples, which creates a severe temporal credit-assignment problem across the denoising process. We propose Latent Reward Registers, a mechanism that estimates terminal preference directly from intermediate noisy latents. Learnable, position-free register tokens are appended as an auxiliary read path to a frozen Diffusion