← Back to all articles
arXiv cs.AIOctober 7, 2026

Language Model Activations Inhabit Privileged Error-Correcting Basins

Excerpt

arXiv:2610.04183v1 Announce Type: new Abstract: Language models exhibit remarkable robustness, continuing to produce coherent text even when their activations are perturbed by interventions like linear steering. We hypothesize that this robustness is a result of passive dynamics, i.e., constraining mechanisms in the forward pass that funnel activations toward "good" regions that produce coherent outputs. To investigate these hypothesized error-correcting mechanisms, we probe the geometry of lang