arXiv cs.LGOctober 2, 2026
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
Excerpt
arXiv:2604.18572v3 Announce Type: replace-cross Abstract: The Platonic Representation Hypothesis posits that neural networks trained on different modalities (e.g., text and images) converge toward a shared representation of reality. If true, this has significant implications for whether modality choice matters at all. In this paper, we show that the evidence for this claim is substantially weaker than subsequent work suggests. The mutual $k$-nearest-neighbor metric used on 1024 text-image pairs