arXiv cs.AIOctober 7, 2026
Backdooring Sparse Autoencoders
Excerpt
arXiv:2610.06049v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) are increasingly used not only to interpret language models but also to intervene on their internal representations. We show that this creates a supply-chain attack surface: a maliciously modified SAE can induce attacker-chosen behavior when inserted into the forward pass of an otherwise unchanged language model. We introduce a decoder-only SAE backdoor that leaves both the underlying LLM and the SAE encoder frozen, res