arXiv cs.CLSeptember 10, 2026
StreamAlign: Streaming Text-Aligned Speech Tokenization
Excerpt
arXiv:2609.09719v1 Announce Type: new Abstract: Text-aligned speech tokenization methods have emerged to better align speech tokens with LLM token spaces, enabling more effective utilization of pretrained LLMs. However, they rely on offline automatic speech recognition (ASR), leading to two key limitations: (i) the need for complete utterances before tokenization, precluding real-time streaming, and (ii) vocabulary mismatch between ASR and LLMs, which reduces acoustic granularity from the subwor