← Back to all articles
arXiv cs.LGAugust 17, 2026

SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos

Excerpt

arXiv:2510.02916v2 Announce Type: replace-cross Abstract: We propose SALSA-V, a multimodal video-to-audio generation model capable of synthesizing highly synchronized, high-fidelity long-form audio from silent video content. Our approach introduces a masked diffusion objective, enabling audio-conditioned generation and the seamless synthesis of unconstrained length audio sequences. Additionally, by integrating a shortcut loss into our training process, we achieve rapid generation of high-quality