← Back to all articles
arXiv cs.LGOctober 1, 2026

Distribution Matching Distillation for Continuous Diffusion Language Models

Excerpt

arXiv:2609.40235v1 Announce Type: new Abstract: Continuous diffusion language models generate all tokens in parallel, yet high-quality generation can still require hundreds of network evaluations (NFEs). We study how distributional distillation can reduce this cost by exploiting the student's probabilistic token outputs. Our unified formulation connects the student's output parameterization to the resulting gradient estimators and yields two methods with the same student architecture and reverse