arXiv cs.LGOctober 7, 2026
MASKerade: Token-Routed Mask Experts for Dense-to-MoE Upcycling
Excerpt
arXiv:2610.07809v1 Announce Type: new Abstract: Sparsely activated Mixture-of-Experts (MoE) models increase model capacity without a proportional increase in per-token computation. Dense-to-MoE upcycling reuses pretrained dense models to construct such systems, commonly by copying feed-forward networks (FFNs) into independently trained experts. We introduce MASKerade, a dense-to-MoE training method that instead learns experts as sparse subnetworks of a frozen pretrained FFN. Each expert is defin