arXiv cs.LGOctober 7, 2026
Best-of-$N$ Guidance for Test-time Diffusion Alignment
Excerpt
arXiv:2610.05108v2 Announce Type: replace Abstract: Diffusion models achieve strong generative performance but often struggle to align generated samples with human preferences measured by a reward model. A simple yet effective algorithm for test-time alignment is Best-of-$N$ (BoN) sampling, which draws $N$ i.i.d. samples from a pre-trained diffusion model and outputs the single highest-reward sample. Despite its empirical success, BoN makes limited use of reward information, as it is incorporate