← Back to all articles
arXiv cs.CLSeptember 22, 2026

Time-Incremental Continued Pretraining of LLMs: Knowledge Updates Without Catastrophic Forgetting

Excerpt

arXiv:2609.23916v1 Announce Type: new Abstract: Large language models (LLMs) drift out of date the moment their pretraining ends, yet retraining from scratch is prohibitively expensive. Continued pretraining (CPT) is the natural remedy, but it is typically evaluated through a continual learning lens that assumes disjoint data streams. This is a poor fit for time-incremental updates on web-scale crawls, where successive snapshots share substantial URL overlap by design. We study time-incremental