← Back to all articles
arXiv cs.AIOctober 7, 2026

A Testable Theory of Atomic Features

Excerpt

arXiv:2610.05794v1 Announce Type: new Abstract: We develop and test a theory of language model representations in which there exist atomic features. Our main theoretical insight is that in such a model, sparse dictionaries (e.g., SAEs) of increasing size recover an increasing prefix of the most prevalent atoms in the training data. This "recovery principle" yields three testable predictions: many features in small SAEs are shared by all larger SAEs, SAEs trained on different data share features