arXiv cs.LGOctober 1, 2026
ReDiF: Resource-Efficient Few-Step Diffusion Distillation via Reinforcement Learning
Excerpt
arXiv:2512.22802v2 Announce Type: replace Abstract: Step distillation accelerates diffusion sampling by training a few-step student to imitate a many-step teacher, but distillation itself remains expensive. Typically, this requires thousands of GPU-hours and a large pre-generated trajectory dataset. We introduce ReDiF, which casts step distillation as terminal-reward policy optimization rather than step-wise regression. The student is optimized against a reward computed on the terminal sample, m