arXiv cs.CLSeptember 24, 2026
Routing-Aware Expert Calibration for Machine Unlearning in Mixture-of-Experts Language Models
Excerpt
arXiv:2606.10338v2 Announce Type: replace Abstract: Machine unlearning is increasingly important for large language models, yet unlearning in Mixture-of-Experts (MoE) architectures remains underexplored. Unlike dense models, MoE architectures employ a router at each layer to assign each token to a sparse subset of experts. In this work, we observe that forget data often activates a small subset of experts disproportionately, while these experts may receive much weaker activation from retain data