← Back to all articles
arXiv cs.LGOctober 1, 2026

VOSSA: Voiceprint Optimization for Streaming Speech Architectures

Excerpt

arXiv:2609.38887v1 Announce Type: cross Abstract: Real-time voice conversion (VC) systems commonly rely on pretrained speaker embeddings from automatic speaker verification (ASV) models. While effective for speaker discrimination, these embeddings are trained to remain stable across phonetic and prosodic variations within-speaker, which may conflict with frame-level acoustic generation in streaming constraints. To address this issue, we propose VOSSA (Voiceprint Optimization for Streaming Speech