arXiv cs.CLSeptember 22, 2026
COT-TTS: Audio Context-Aware Text-to-Speech with Chain-of-Thought Reasoning
Excerpt
arXiv:2609.22697v1 Announce Type: new Abstract: Recently, text-to-speech systems have made significant progress in speech expressiveness and controllability. However, the speaking style of generated speech typically relies on clear user-specified instructions. In natural conversations, speaking style should be naturally inferred from the preceding conversational context. Therefore, we propose COT-TTS, a context-aware, reasoning-based text-to-speech task. Given historical conversation audio, targ