arXiv cs.LGOctober 7, 2026
Enhancing Diffusion Language Models with Autoregressive Post-Training Weights
Excerpt
arXiv:2610.08108v1 Announce Type: new Abstract: Diffusion language models (dLLMs) have emerged as a promising alternative to autoregressive (AR) language models, offering flexible token-update orders and parallel decoding. Recent dLLMs are often initialized from pretrained AR models before diffusion conversion in order to inherit their learned representations. After the conversion, however, they typically ignore the extensive post-training ecosystem of their AR ancestors. In this work, we show t