arXiv cs.LGOctober 1, 2026
Contrastive Representation Shaping for LLM Unlearning
Excerpt
arXiv:2601.22028v2 Announce Type: replace Abstract: Most LLM unlearning methods aim to approximate retrain-from-scratch behaviors with minimal distribution shift, often via alignment-style objectives defined in the prediction space. While effective at reducing forgotten content generation, such approaches may act as suppression: forgotten concepts can persist in representations and remain entangled with retained knowledge. We introduce CLReg, a contrastive representation regularizer that identif