arXiv cs.CLSeptember 21, 2026
SG-Mamba: Sparse Graph-Guided Mamba for Audio-Visual Speech Enhancement
Excerpt
arXiv:2609.18009v2 Announce Type: replace-cross Abstract: Lightweight audio-visual speech enhancement (AVSE) models face a critical trade-off between computational efficiency and cross-modal alignment accuracy. While simple concatenation lacks relational expressiveness, dense cross-attention incurs computational overhead and is prone to unreliable cross-modal correspondence under strong acoustic interference. We propose Sparse Graph-Guided Mamba (SG-Mamba), a lightweight AVSE framework that inte