arXiv cs.AIOctober 7, 2026
Off-Policy Merging Beats On-Policy Self-Distillation for Continual Learning
Excerpt
arXiv:2610.05872v1 Announce Type: cross Abstract: A long-standing goal of AI is a model that can continually learn and improve itself. On post-trained models, supervised finetuning (SFT) on new data often causes poor generalization and catastrophic forgetting. As such, the conventional wisdom is that on-policy training is a prerequisite for continual learning. In practice, however, data containing new knowledge or capabilities are often off-policy. While methods such as on-policy self-distillati