arXiv cs.CLOctober 7, 2026
Quantifying the Generation Modality Gap in Speech-Text Language Models
Excerpt
arXiv:2609.14743v2 Announce Type: replace Abstract: Pure speech language models often lag behind text and speech-text language models in generating coherent content, but this gap is difficult to quantify because speech and text systems are typically evaluated with different metrics and trained on different data. We study the speech-text modality gap in a family of spoken language models, based on flow matching for continuous acoustic feature generation. We construct a unified generation-based ev