arXiv cs.LGOctober 2, 2026
Embedding Prediction Helps Image Generation
Excerpt
arXiv:2610.02203v1 Announce Type: cross Abstract: In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition