← Back to all articles
arXiv cs.CLSeptember 14, 2026

SynthSentry: Detecting Synthetic Data Contamination in Language Model Training Data

Excerpt

arXiv:2609.12353v1 Announce Type: new Abstract: Large language models trained recursively on their own or other models' outputs undergo model collapse, in which distributional tails and factual accuracy deteriorate while fluency survives. Prior work diagnoses collapse after training; the actionable problem is screening a corpus of unknown provenance before training. We introduce SynthSentry, a corpus-level, model-agnostic contamination signal requiring no access to the generating model, no gener