arXiv cs.AIOctober 7, 2026
Anchored or Drifting: What Recursive Self-Generation Reveals About Training Data
Excerpt
arXiv:2606.31991v2 Announce Type: replace-cross Abstract: Large generative models are known to memorize their training data, posing severe privacy risks. Yet, current methods to detect training membership typically rely on the weak signals of a single forward pass. In this work, we find that training samples and unseen (held-out) data follow visibly different trajectories under recursive self-generation -- repeatedly feeding a model's output back as its next input. Held-out samples \emph{drift}: