arXiv cs.AIOctober 7, 2026
On the Geometry of Multimodal Saturation: Riemannian VICReg
Excerpt
arXiv:2610.06096v1 Announce Type: cross Abstract: In self-supervised learning, a third modality should improve, or at least preserve, performance. Across nine image-text-tabular datasets, we show that it instead harms performance: the trimodal model underperforms its own best bimodal subset in 55.6% of paired runs under VICReg. The same failure occurs in 51.1% of paired runs under SimSiam. We call this failure multimodal saturation. We propose that the failure lies in the alignment geometry. Rie