arXiv cs.CLSeptember 18, 2026
To Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals
Excerpt
arXiv:2609.20186v1 Announce Type: new Abstract: Speculative Decoding (SD) has significantly accelerated Large Language Model (LLM) inference, yet existing approaches face a fundamental tradeoff between two drafting strategies: neural drafting and context-based copying. Neural drafts (e.g., EAGLE3) provide robust performance across diverse text settings, while copy-based methods achieve higher speedups in copy-intensive regimes by generating candidates faster and exploiting long repetition spans