arXiv cs.LGOctober 7, 2026
Finding the Heads and the Neurons Responsible for Network Information Retrieval in Language Models
Excerpt
arXiv:2610.08200v1 Announce Type: new Abstract: We ask whether specific attention heads, and more finely specific neurons inside those heads, are responsible for recognizing that a language model's context contains network infrastructure information (a hostname paired with its IP address), and whether that responsibility can be validated causally rather than by correlation alone. At the head level the answer is yes, across five models spanning three architecture families: in every model, a small