arXiv cs.CLAugust 19, 2026
SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization
Excerpt
arXiv:2608.13538v2 Announce Type: replace Abstract: Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features still relies primarily on external observation. This reliance leads to superficial explanations inferred from observed model behavior and computational inefficiency from collecting such behavioral evidence at scale. We introduce SAEVerbalizer, a framework that injects SAE decoder directions into an