← Back to all articles
arXiv cs.LGOctober 2, 2026

Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution

Excerpt

arXiv:2610.00417v1 Announce Type: new Abstract: Repeated training on model-generated data can degrade later models. One possible response is to use provenance when deciding which generated examples to reuse. We test both how reliably that provenance can be recovered and whether it helps identify better training data. Using financial-risk text, we first identify the source of generated passages and then repeat the test after rewriting them. Generator attribution is 98.7% accurate on the original