← Back to all articles
arXiv cs.AIAugust 18, 2026

Broken Symmetry in LLM Refusal: Answer Release Is More Local Than Refusal Restoration

Excerpt

arXiv:2608.15772v1 Announce Type: new Abstract: When a language model refuses to answer a prompt, it is unclear whether the correct answer is erased from its internal representations, or merely suppressed at the output layer. We investigate this mechanism using a controlled withhold setting, which yields perfectly matched answering and refusal trajectories for bidirectional activation patching. We uncover a causal asymmetry in intervention locality under matched causal interventions, which we te