← Back to all articles
arXiv cs.AIOctober 7, 2026

A Step Towards Forgetting: Optimiser History and the Loss of Answer Mass

Excerpt

arXiv:2610.03940v1 Announce Type: cross Abstract: During fine-tuning, a language model can assign less probability to previously learned answers even when the current gradient acts to preserve that probability. With momentum, each update also carries gradients computed at earlier model states, and these stored contributions can push the model in the opposite direction. We investigate how this optimiser memory contributes to forgetting by separating old-task loss into confusion among its answers