arXiv cs.LGOctober 1, 2026
Fork-dLLM: Avoiding the Flexibility Trap in Diffusion Language Models
Excerpt
arXiv:2609.39859v1 Announce Type: new Abstract: Masked diffusion language models (dLLMs) have shown strong potential for faster inference through parallel token generation when combined with confidence-based samplers. However, recent work has shown that such methods can defer unmasking high-entropy fork positions at which multiple plausible continuations exist. This results in reduced generation diversity, as shown by worse pass@k scaling, and limits gains obtainable from RL post-training. To av