arXiv cs.CLSeptember 24, 2026
Does Step Law Transfer to Small-Scale Language Models? An Empirical Recalibration Below 59M Parameters
Excerpt
arXiv:2609.27581v1 Announce Type: cross Abstract: Step Law gives power-law formulas for the optimal peak learning rate eta* and batch size B* when pre-training language models. It was calibrated on models between 59M and 1B parameters; the small-model regime N < 59M was never tested empirically by its authors. This regime matters for single-GPU training, interpretability research, educational experiments, and settings where larger models are infeasible on memory or cost grounds. We test whether