arXiv cs.AIOctober 7, 2026
ScopeSAE: Model-Scope Feature Discovery with Interpretable Layer Selection
Excerpt
arXiv:2610.04905v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) are a central tool in mechanistic interpretability. However, existing SAEs are primarily trained per layer. The modeling subspace is therefore fixed by layer identity, independent of which token-layer states actually drive each prediction. We argue that this constraint contributes to several limitations observed in layer-wise SAEs, including low feature utilization, high dictionary redundancy, and features that lack dir