arXiv cs.CLOctober 7, 2026
Does Steering Break Your Model? A Multi-Dimensional Evaluation Suite for LLM Steering Methods
Excerpt
arXiv:2610.07722v1 Announce Type: new Abstract: Activation steering provides a lightweight and flexible way to control large language model (LLM) behavior. However, effective steering requires more than inducing the intended behavior: it should also limit unintended changes and remain robust across inputs and training data. Existing evaluations cover these dimensions only in fragments. As a result, the trade-offs between efficacy and side effects have not been systematically characterized. We in