← Back to all articles
arXiv cs.LGOctober 2, 2026

Match the Distribution, Not the Compute: Post-Training Multi-Token Prediction Heads

Excerpt

arXiv:2610.00888v1 Announce Type: new Abstract: Multi-token prediction (MTP) improves the throughput of autoregressive generation by enabling the language model to draft multiple next tokens per forward pass, while a verification step over draft tokens ensures that token distribution of the backbone is preserved. Every open MTP-family release (MiMo-7B, DeepSeek-V3, Qwen3) trains its heads jointly with the backbone over the full pretraining run of tens of trillions of tokens, thus setting the dra