arXiv cs.LGOctober 2, 2026
Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution
Excerpt
arXiv:2610.00417v1 Announce Type: new Abstract: Repeated training on model-generated data can degrade later models. One possible response is to use provenance when deciding which generated examples to reuse. We test both how reliably that provenance can be recovered and whether it helps identify better training data. Using financial-risk text, we first identify the source of generated passages and then repeat the test after rewriting them. Generator attribution is 98.7% accurate on the original