arXiv cs.LGAugust 18, 2026
Domain-Agnostic Neural Topic Modeling with Contextual Token-Level Semantic Graph Representation
Excerpt
arXiv:2608.16269v1 Announce Type: cross Abstract: Recent advances in neural topic models with pre-trained language models (PLMs) have achieved strong performance by leveraging general-domain pre-training, yet their topic interpretability often degrades on specialized corpora. This limitation primarily stems from the geometry of the embedding space, where domain-specific terms unseen during pre-training collapse into an indistinguishable region, and neither domain-specific re-training, word-level