arXiv cs.LGOctober 2, 2026
Universal interpolation for deep residual self-attention networks
Excerpt
arXiv:2610.01981v1 Announce Type: new Abstract: Universal approximation is a necessary qualitative property of learning architectures to benefit from scaling laws. While it is generically verified on a variety of neural architectures and random feature models, it typically involves infinite width limits. In this work, we focus on deep self-attention models and consider instead the `dual' regime, where approximation power is enabled entirely by depth, and featuring strong parameter sharing across