arXiv cs.LGOctober 2, 2026
Match the Distribution, Not the Compute: Post-Training Multi-Token Prediction Heads
Excerpt
arXiv:2610.00888v1 Announce Type: new Abstract: Multi-token prediction (MTP) improves the throughput of autoregressive generation by enabling the language model to draft multiple next tokens per forward pass, while a verification step over draft tokens ensures that token distribution of the backbone is preserved. Every open MTP-family release (MiMo-7B, DeepSeek-V3, Qwen3) trains its heads jointly with the backbone over the full pretraining run of tens of trillions of tokens, thus setting the dra