← Back to all articles
arXiv cs.LGOctober 7, 2026

SpecFold: Folding Multi-Branch Redundancy for Faster Speculative Decoding in Diffusion Language Models

Excerpt

arXiv:2610.04875v2 Announce Type: replace-cross Abstract: Diffusion large language models (DLLMs) generate text through iterative block denoising, and multi-branch speculative decoding accelerates this process by verifying a main branch together with multiple draft branches in a single forward pass. While prior DLLM acceleration methods primarily exploit temporal redundancy across denoising steps, we identify a complementary redundancy axis within each speculative verification step: multi-branch