← Back to all articles
arXiv cs.CLSeptember 18, 2026

Causal Analysis and Mitigation of Spurious Onsets in Full-Duplex Speech LLMs

Excerpt

arXiv:2609.13445v2 Announce Type: replace Abstract: Speech-to-speech LLMs like Moshi, and its derivative PersonaPlex, can listen and speak concurrently through full-duplex generation. However, they can begin speaking inappropriately during prolonged user silence: under digital-zero input, Moshi and PersonaPlex initiate speech in 30% and 27.5% of five-minute continuations, respectively. What causes this spurious speech? We investigate two hypotheses: either repeated sampling selects speech despit