arXiv cs.LGOctober 7, 2026
Ideal Paths for Approximating Logistic Gradient Descent Trajectories at Large Initialization
Excerpt
arXiv:2610.04142v2 Announce Type: replace Abstract: Modern training on a new task often starts from a previously trained model rather than from scratch, raising the question of how this initialization affects the subsequent training trajectory. Classical implicit-bias results characterize the direction selected by prolonged training, but this direction alone does not provide information regarding the intermediate behavior. We address this question through a geometric approximation of full-batch