← Back to all articles
arXiv cs.LGOctober 1, 2026

Emergent alignment and the projectability of ethical personas

Excerpt

arXiv:2606.09475v3 Announce Type: replace-cross Abstract: Work on `emergent misalignment' shows that finetuning LLMs on narrow tasks can induce broadly misaligned behavior. This supports the `persona selection' (PSM) hypothesis: during pre-training, LLMs learn to simulate different characters/perspectives, which can be elicited and refined during post-training. This paper investigates the converse phenomenon, `emergent alignment', and uses it to support and refine the PSM and motivate a novel de