← Back to all articles
arXiv cs.CLSeptember 22, 2026

Structure Before Sampling: Community-Aware Core-Set Selection for Data-Efficient Text-to-Speech

Excerpt

arXiv:2609.24275v1 Announce Type: new Abstract: Text-to-speech (TTS) corpora are costly to record, yet many utterances add little new phonetic information. Core-set selection reduces this cost by choosing a small training subset under a fixed audio-duration budget. We represent a corpus as a phonotactic graph that links each utterance to its most phonemically similar ones, and we first test whether this graph has structure. In Bangla and English corpora, its clustering is 199 and 56 times that o