← Back to all articles
arXiv cs.LGOctober 1, 2026

LLM Persona Unlearning

Excerpt

arXiv:2609.39882v1 Announce Type: cross Abstract: Pre-training equips large language models (LLMs) with a broad repertoire of behavioral patterns associated with roles, styles, values, and goals. Post-training teaches conditional enactment and makes a helpful Assistant the default, but it does not erase alternative modes from the weights; explicit prompts can therefore elicit personas that repeatedly shape judgment, language, and action. In open-weight settings, runtime controls can be removed,