arXiv cs.AIOctober 7, 2026
Weight Decay and Neuron Condensation: A Three-Stage Analysis of Two-Layer ReLU Networks
Excerpt
arXiv:2610.04533v1 Announce Type: cross Abstract: Weight decay is widely used as a regularization technique in neural network training, yet its role in neuron condensation (parameter direction alignment) remains unclear. Starting from a parameter initialization in the neural tangent kernel regime, we characterize training dynamics under weight decay through three stages: rapid fitting, amplitude compression, and neuron condensation. Using a two-layer ReLU network, we analyze a residual correlati