arXiv cs.AIOctober 2, 2026
Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness
Excerpt
arXiv:2610.00568v1 Announce Type: cross Abstract: Large language models are characterized by three key properties: capability, alignment, and faithfulness. Prior work studies the tradeoffs between capability and alignment, and between capability and faithfulness, but a third tension remains underexplored: the alignment-faithfulness conflict. We show that aligned models systematically deviate from their inputs on unsafe or sensitive content without disclosing the modification, a failure mode we c