arXiv cs.CLSeptember 11, 2026
Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data
Excerpt
arXiv:2609.11917v1 Announce Type: cross Abstract: As the supply of human-written text is exhausted, it has become standard practice to repeat language model training data. Prior work has studied data repetition for densely activated Transformers, but the effects of data repetition remains largely unexplored for recently dominant sparse architectures such as Mixture-of-Experts (MoE), despite their increased compute efficiency. We vary data repetition rates across single- and multi-domain data mix