← Back to all articles
arXiv cs.AIOctober 7, 2026

Routing-Aware Safety Alignment for Mixture-of-Experts Models

Excerpt

arXiv:2602.04448v3 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) language models introduce unique challenges for safety alignment due to their sparse routing mechanisms, which can enable degenerate optimization behaviors under standard full-parameter fine-tuning. In our preliminary experiments, we observe that naively applying full-parameter safety fine-tuning to MoE models can reduce attack success rates through routing or expert dominance effects, rather than by directly repa