arXiv cs.LGOctober 1, 2026
How Does Local Landscape Geometry Evolve in Language Model Pre-Training?
Excerpt
arXiv:2609.39767v1 Announce Type: new Abstract: The scale and expense of pre-training language models make efficient hyperparameter tuning essential, yet a principled guidance is still missing. In this work, we analyze language model pre-training dynamics from a local landscape geometry perspective. Our study reveals two distinct phases. In Phase I, sharpness of the local landscape is initially high, leading to instability and loss plateaus under large learning rates (LRs). The landscape shifts