arXiv cs.LGOctober 2, 2026
Do Vision Transformers Need All-to-All Attention? Global Communication Through Elastic Learned Cores
Excerpt
arXiv:2605.12491v2 Announce Type: replace-cross Abstract: Vision Transformers (ViTs) learn rich visual-semantic representations through all-to-all self-attention among patch tokens. However, this design implicitly assumes that direct pairwise patch interactions are necessary for effective representation learning. In this work, we challenge this assumption and show that representations supporting both global recognition and dense prediction can be learned without direct patch-to-patch interaction