arXiv cs.CLSeptember 10, 2026
Scaling phoneme-based TTS augmentation for ASR: A unified pipeline and controlled study
Excerpt
arXiv:2608.26697v2 Announce Type: replace Abstract: Synthetic speech offers scalable supervision for automatic speech recognition (ASR), but its benefit depends on text selection, reference speech, and augmentation scale. We present a phoneme-based TTS-to-ASR pipeline using a single TTS model jointly trained from scratch on Arabic, French, Italian, and Portuguese with the F5-TTS architecture and language-monolingual ASR systems cover 13 test sets. Across the synthesis-scale sweep, random augment