← Back to all articles
arXiv cs.CLSeptember 23, 2026

PACE-dLLM: Elastic Block Decoding via Confidence Cliff Estimation for Diffusion Language Models

Excerpt

arXiv:2609.26249v1 Announce Type: new Abstract: Diffusion language models (dLLMs), such as LLaDA and Dream, have become competitive with autoregressive (AR) LLMs in generation quality while supporting native parallel decoding. A standard acceleration strategy is block-wise decoding, where each forward pass predicts a block of length B and commits high-confidence tokens. However, B couples two distinct decisions: the look-ahead horizon and the number of tokens to commit. Existing accelerators add