arXiv cs.LGOctober 1, 2026
From Spectra to Joint Schedules in LLM Pre-training: 3+3(+2) Scaling-Law Regimes
Excerpt
arXiv:2609.40148v1 Announce Type: new Abstract: Power-law learning curves are often treated as fixed properties of a model and its data, although learning-rate and batch-size schedules can change the observed loss. We study this dependence in noisy online SGD with linear random features. Conditional on the representation, an exact Volterra equation separates two response components: a forcing term that propagates unresolved target error and a memory kernel that propagates stochastic-error inject