← Back to all articles
arXiv cs.LGOctober 2, 2026

Forking: Sudden Overfitting Under Replay

Excerpt

arXiv:2610.00394v1 Announce Type: new Abstract: This paper studies forking, a generalization failure discovered in NanoGPT autoresearch. Under data replay, models with an over-encoding n-gram memory branch show a sharp separation of training and validation loss at epoch boundaries, resembling the shape of forks. We study this phenomenon in a controlled vanilla NanoGPT setting and reproduce it in a DeepSeek-style model with Engram. Mechanistically, repeated updates sharpen the continuations obser