arXiv cs.CLSeptember 23, 2026
Recovering the Zipfian Distribution in Unsupervised Term Discovery
Excerpt
arXiv:2606.10781v3 Announce Type: replace-cross Abstract: Unsupervised term discovery involves segmenting unlabelled speech into word- or syllable-like units and clustering these into a lexicon of candidate types. True lexicons follow a Zipfian distribution, yet the dominant centre-based clustering approach -- K-means -- produces a more uniform distribution due to an inductive bias toward spherical clusters. In this paper we revisit graph-based clustering as a bottom-up alternative, where segmen