← Back to all articles
arXiv cs.CLSeptember 22, 2026

HaikuS2S: A Cascaded System For Responding In Verse

Excerpt

arXiv:2609.23951v1 Announce Type: new Abstract: Expressive speech synthesis has advanced through prosody modeling, yet generating structured poetic speech, such as haiku, remains challenging. Prior work on prosody transfer improves expressiveness, and fine-tuned poetry TTS (text-to-speech) systems capture verse intonation. However, these models do not model haiku's 5-7-5 syllable structure or line-ending pauses. We present a cascaded system, HaikuS2S, combining ASR (automatic speech recognition)