arXiv cs.CLSeptember 22, 2026
Structure Before Sampling: Community-Aware Core-Set Selection for Data-Efficient Text-to-Speech
Excerpt
arXiv:2609.24275v1 Announce Type: new Abstract: Text-to-speech (TTS) corpora are costly to record, yet many utterances add little new phonetic information. Core-set selection reduces this cost by choosing a small training subset under a fixed audio-duration budget. We represent a corpus as a phonotactic graph that links each utterance to its most phonemically similar ones, and we first test whether this graph has structure. In Bangla and English corpora, its clustering is 199 and 56 times that o