← Back to all articles
arXiv cs.AIOctober 2, 2026

DRelay: Global Draft Context for Prefix-Aware Parallel Speculative Decoding Repair

Excerpt

arXiv:2610.01439v1 Announce Type: new Abstract: Parallel drafting reduces the drafting overhead of speculative decoding for large language models (LLMs), but its gains remain limited by the accepted prefix length. Even when the correct token is present in the candidate pool, a single early selection error prevents subsequent predictions from being used. We propose DRelay, which uses global information from the entire draft block to perform prefix-aware selective repair of candidate selections be