arXiv cs.AIOctober 7, 2026
SharpDraft: Accelerating Long-Context Speculative Decoding with Cardinality-Aware Query Scaling
Excerpt
arXiv:2610.05106v1 Announce Type: new Abstract: Long-form reasoning makes inference expensive, and speculative decoding mitigates this cost by verifying multiple draft tokens in parallel. Its speedup, however, can fade as context grows and draft acceptance declines. We focus on attention-mass dilution: as softmax normalizes over more visible Keys, the mass concentrated on the highest-scoring Keys can decrease. We introduce SharpDraft, a training-free method that counteracts this effect through c