arXiv cs.CLSeptember 11, 2026
OUTLETS: Output-Length Prediction from Speculative Decoding Backbones
Excerpt
arXiv:2609.01068v2 Announce Type: replace Abstract: The heavy-tailed distribution of output lengths in Large Language Model (LLM) serving poses major challenges for resource provisioning and cluster scheduling. Although output-length prediction can mitigate these issues, existing approaches have key drawbacks: external proxy models add substantial latency and often have limited fidelity, whereas internal state-based methods are efficient but rely on shallow probes of current model states. We ide