← Back to all articles
arXiv cs.CLSeptember 28, 2026

Learning Natural Conversational Behavior in Tandem Speech-to-Speech Models with Randomized Guidance

Excerpt

arXiv:2609.30773v1 Announce Type: new Abstract: Tandem speech-to-speech architectures couple a responsive speech frontend with an asynchronous text backend. In KAME, a large language model (LLM) serves as the backend, supplying candidate responses as guidance to the speech frontend while the user is still speaking. Ordinary conversation recordings capture the eventual response but not the guidance the backend would supply during the user's utterance. Generating the missing guidance with a simula