← Back to all articles
arXiv cs.LGOctober 1, 2026

Recovering Off-Policy Supervision for Speculative Decoding

Excerpt

arXiv:2609.38795v1 Announce Type: cross Abstract: Block drafters for speculative decoding are commonly trained on corpora written by external models, where a single off-policy token invalidates supervision for all subsequent slots in a block. Existing approaches discard these divergent slots, resulting in severe supervision loss. To resolve this problem while preserving the training corpus, we propose a rollout-based training framework that recovers full supervision through two complementary com