arXiv cs.CLSeptember 11, 2026
A Fragility Spectrum for Recursive Language-Model Training
Excerpt
arXiv:2609.11149v1 Announce Type: new Abstract: Model-generated text is finding its way back into training corpora, and there is plenty of evidence that training on such data over and over collapses output diversity. Prior work has studied the phenomenon itself: which protocols and which data mixtures cause collapse. But different models behave very differently under the same process. We fix one recursive contamination protocol and let 13 publicly released checkpoints form an ecosystem that shar