arXiv cs.LGOctober 2, 2026
Structured-Noise Masked Modeling for Video, Audio and Beyond
Excerpt
arXiv:2503.16311v2 Announce Type: replace Abstract: Masked modeling has emerged as a robust self-supervised learning framework. However, most methods rely on random masking, which disregards the structural properties of different data modalities. To align with the spatiotemporal and spectral characteristics of video and audio data, we introduce a structured noise-based masking approach. By filtering white noise into different color noise distributions, we generate structured masks that capture m