← Back to all articles
arXiv cs.LGOctober 1, 2026

How Does Local Landscape Geometry Evolve in Language Model Pre-Training?

Excerpt

arXiv:2609.39767v1 Announce Type: new Abstract: The scale and expense of pre-training language models make efficient hyperparameter tuning essential, yet a principled guidance is still missing. In this work, we analyze language model pre-training dynamics from a local landscape geometry perspective. Our study reveals two distinct phases. In Phase I, sharpness of the local landscape is initially high, leading to instability and loss plateaus under large learning rates (LRs). The landscape shifts