arXiv cs.LGOctober 1, 2026
Aligned Data Can Induce Misalignment via Context Confusion
Excerpt
arXiv:2609.38379v1 Announce Type: cross Abstract: Large language models (LLMs) are frequently updated for various use cases, where filtering out misaligned training samples is a common practice for preventing post-update misalignment. However, alignment is inherently context-dependent: a recommendation that is aligned in one context may be inappropriate in another. For example, in response to the question "What should a researcher do with the research data?", recommending that the researcher pre