arXiv cs.LGOctober 1, 2026
Right Answer, Wrong Mechanism: Detecting Pernicious Divergence in Causal Interventions
Excerpt
arXiv:2609.39243v1 Announce Type: new Abstract: Causal interventions such as activation patching and distributed alignment search (DAS) are the main tool for making mechanistic claims about neural networks. Recent work showed that these interventions routinely push representations off the model's natural distribution, and that such divergence is sometimes harmless and sometimes pernicious: it can recruit pathways the model never uses on natural inputs, so that an intervention produces the expect