arXiv cs.AIOctober 2, 2026
DRelay: Global Draft Context for Prefix-Aware Parallel Speculative Decoding Repair
Excerpt
arXiv:2610.01439v1 Announce Type: new Abstract: Parallel drafting reduces the drafting overhead of speculative decoding for large language models (LLMs), but its gains remain limited by the accepted prefix length. Even when the correct token is present in the candidate pool, a single early selection error prevents subsequent predictions from being used. We propose DRelay, which uses global information from the entire draft block to perform prefix-aware selective repair of candidate selections be