← Back to all articles
arXiv cs.LGOctober 2, 2026

CAST: Cost-Aware Speculative Trees from One-Pass Block Drafters

Excerpt

arXiv:2610.00321v1 Announce Type: cross Abstract: Speculative decoding accelerates large language model inference by drafting future tokens cheaply and verifying them with the target model in parallel. Block drafters score a whole block of future tokens in one forward pass, yet standard decoding verifies only the top-scoring chain and discards the other candidates. Because these candidates are already scored, verifying more of them adds target computation but no extra drafting. We introduce CAST