← Back to all articles
arXiv cs.AIOctober 2, 2026

AVSD-Scenes: A Dataset for Audio-Visual Description of Urban Scenes

Excerpt

arXiv:2610.01861v1 Announce Type: new Abstract: Natural language descriptions can provide rich semantic representations of audio-visual urban scenes, yet datasets that jointly describe both auditory and visual information remain limited. In this paper, we introduce AVSD-Scenes, a paired audio-visual scene description dataset for urban environments. The dataset contains 12,291 audio-visual scene descriptions generated from the TAU Urban Audio-Visual Scenes dataset. To construct the dataset, we fi