← Back to all articles
arXiv cs.LGOctober 7, 2026

Component and Dimension Sparsity in Transformer Refusal Mechanisms

Excerpt

arXiv:2610.06903v1 Announce Type: cross Abstract: Activation steering manipulates large language model behavior by intervening on internal activations, but the mechanistic basis of these interventions remains poorly understood. We decompose refusal steering into component-level interventions across four open-weight models, identifying the sparse subsets of attention and MLP components whose steering suffices to reproduce the full behavioral effect. We find that refusal directions concentrate in