arXiv cs.CLOctober 7, 2026
DLoop: Looped Speculative Decoding
Excerpt
arXiv:2610.07659v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive generation in large language models. In each drafting stage, a lightweight draft model proposes tokens that the target model subsequently verifies. With increasingly capable draft models, we find that the target model frequently accepts all tokens produced in a drafting stage. A verification nevertheless follows each drafting stage, resulting in unnecessary target-model forward passes even when drafti