← Back to all articles
arXiv cs.AIAugust 17, 2026

Parameter- and Bandwidth-Efficient Edge--cloud Many-to-Many Speech-to-Text Translation

Excerpt

arXiv:2605.28642v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have demonstrated significant potential for speech-to-text translation (S2TT). However, existing deployment paradigms face critical challenges: pure on-device models suffer from resource constraints, while centralized cloud systems incur bandwidth bottlenecks and privacy risks by transmitting raw voice data. In this paper, we propose Edge--cloud Speech Recognition and Translation (ESRT), a parameter-effi