← Back to all articles
arXiv cs.LGOctober 1, 2026

Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior

Excerpt

arXiv:2609.39827v1 Announce Type: cross Abstract: Pre-pretraining (PPT) on synthetic non-natural language data improves token efficiency during language model pre-training (PT). Prior work attributes this gain to a grammatical prior, i.e., a structural inductive bias learned during PPT that transfers to natural language grammar. However, PPT has only been tested on models of at most 1B parameters and PT budgets below 2B tokens on predominantly web text. It is unknown whether PPT is effective at