← Back to all articles
arXiv cs.LGOctober 1, 2026

Learning What to Forget: Distributional Unlearning for LLM Representation Spaces

Excerpt

arXiv:2609.38929v1 Announce Type: cross Abstract: Machine learning systems increasingly face the need to remove the influence of entire data domains, such as toxic language, harmful behavior, or topical content, rather than isolated records. Recent work formalizes this problem as \emph{distributional unlearning}: selecting a subset of a forget domain whose removal moves the training distribution away from an unwanted population while preserving proximity to the desired one. However, existing ana