← Back to all articles
arXiv cs.LGOctober 2, 2026

Persistent Depth Ordering amid Shifting Block-Bypass Responses in Language Model Pretraining

Excerpt

arXiv:2610.01165v1 Announce Type: new Abstract: Layer interventions are widely used to probe the internal organization of language models, yet most analyses examine a single training checkpoint even though model representations and computations evolve throughout pretraining. This leaves open which depth-dependent intervention responses reflect persistent organization and which are transient consequences of training. We study this question using single-block identity bypass on fixed teacher-force