← Back to all articles
arXiv cs.CLSeptember 11, 2026

Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models

Excerpt

arXiv:2609.10830v1 Announce Type: new Abstract: When a language model finds a sentence unusually cheap to predict, it is tempting to conclude that the sentence was in its training data. Almost every published test of that inference has had to guess which sentences were in the training data, the members, and which were not. This paper removes the guessing. Two model families, OLMo-2 and Pythia, publish their pretraining corpora, and a public index over those corpora returns the exact number of ti