arXiv cs.AIOctober 7, 2026
A Testable Theory of Atomic Features
Excerpt
arXiv:2610.05794v1 Announce Type: new Abstract: We develop and test a theory of language model representations in which there exist atomic features. Our main theoretical insight is that in such a model, sparse dictionaries (e.g., SAEs) of increasing size recover an increasing prefix of the most prevalent atoms in the training data. This "recovery principle" yields three testable predictions: many features in small SAEs are shared by all larger SAEs, SAEs trained on different data share features