arXiv cs.LGOctober 7, 2026
Should We Skip Diffusion?
Excerpt
arXiv:2610.07002v1 Announce Type: new Abstract: Diffusion models learn semantic representations while generating images. In the Decoupled Diffusion Transformer (DDT), a condition encoder provides features that guide a velocity decoder in denoising. To enable effective denoising at all noise levels, these features must capture both high-level abstract structures and low-level details. However, skip/residual connections in the encoder allow shallow features to bypass successive transformations, wh