arXiv cs.LGOctober 1, 2026
Awakening of the Buddha: Subspace Learning During Population-Loss Plateaus
Excerpt
arXiv:2609.39408v1 Announce Type: new Abstract: Population loss can remain nearly constant while a neural network learns a substantially more predictive representation. We establish this separation for two-layer ReLU and leaky-ReLU networks trained on Gaussian inputs by simultaneous fixed-step population gradient descent on all parameters. For structured additive teachers whose links are positive mixtures of Gaussian-damped cubics in $H^1(\gamma)$, we give explicit conditions under which small I