← Back to all articles
arXiv cs.CLSeptember 10, 2026

Beyond Top Words: MonoTM for Topic Modeling with Interpretable Monosemantic Features

Excerpt

arXiv:2609.09575v1 Announce Type: new Abstract: Topic models summarize large text corpora, but top-ranked words often provide only a limited representation of topic semantics. Sparse autoencoders (SAEs) offer a way to move beyond word-level descriptors by extracting interpretable features from dense representations, yet how feature interpretability relates to topic-inference quality remains unclear. We introduce \textbf{MonoTM}, an interpretable topic modeling framework that decouples these role