← Back to all articles
arXiv cs.AIOctober 7, 2026

On the Steering Dimensionality of Refusal in Language Models

Excerpt

arXiv:2610.04245v1 Announce Type: new Abstract: Existing activation steering methods often assume that a high-level concept can be mediated by a single steering direction. To support this, two complementary interventions should be achieved: additive steering should induce the target behavior, while directional ablation should suppress it. Yet behaviors may occupy richer activation geometries beyond a single direction, and semantically similar behaviors may be represented by distinct directions.