← Back to all articles
arXiv cs.LGOctober 7, 2026

Do Sparse Autoencoders Learn Meaningful Concept Hierarchies?

Excerpt

arXiv:2606.22994v2 Announce Type: replace Abstract: Sparse autoencoders (SAEs) have become an important tool for unsupervised concept discovery in large models. To make the resulting feature spaces more interpretable and manageable, recent approaches have begun imposing hierarchical structure, either explicitly or as an implicit effect of training constraints, yet rigorous comparison remains difficult. There are no agreed-upon requirements for what a meaningful feature hierarchy should satisfy,