← Back to all articles
arXiv cs.AIOctober 7, 2026

Slaying the Hydra: Interaction-Aware Circuit Discovery in Language Models

Excerpt

arXiv:2610.04017v1 Announce Type: cross Abstract: Localizing behavior to individual components of a language model is a central goal of mechanistic interpretability. However, scoring components one at a time misses context-dependent effects: a primary component can inhibit the activation of a backup, leading to issues with ranking components. Actual causality studies the structure of such interactions via witnesses: variables that provide contextual information to resolve interaction terms. Howe