arXiv cs.CLSeptember 10, 2026
Do speech foundation models really learn words?
Excerpt
arXiv:2609.10434v1 Announce Type: new Abstract: Self-supervised speech foundation models are now used in a wide array of downstream applications, including traditional speech recognition and as the basis for tokens in speech-aware language models. Attempts to understand their usefulness have largely focused on probing their representations' ability to discriminate phonemes and words. However, discriminative ability for words need not imply specialized representation of words per se. Good discrim