arXiv cs.LGOctober 1, 2026
Provable Benefit of SignGD: A Minimal Model Under Heavy-Tailed Class Imbalance
Excerpt
arXiv:2512.00763v2 Announce Type: replace Abstract: Adaptive and non-Euclidean optimizers often outperform Euclidean methods such as stochastic gradient descent (SGD) in language modeling by a large margin. Existing theory usually explains this gap by assuming favorable smoothness geometry or noise structure tailored to the specific optimizer. We instead ask whether such geometry can be induced from a concrete learning setting. Starting from an optimizer gap that persists across realistic langua