arXiv cs.AIOctober 7, 2026
On the Steering Dimensionality of Refusal in Language Models
Excerpt
arXiv:2610.04245v1 Announce Type: new Abstract: Existing activation steering methods often assume that a high-level concept can be mediated by a single steering direction. To support this, two complementary interventions should be achieved: additive steering should induce the target behavior, while directional ablation should suppress it. Yet behaviors may occupy richer activation geometries beyond a single direction, and semantically similar behaviors may be represented by distinct directions.