arXiv cs.LGOctober 2, 2026
Pre-training interventions, ex post facto: Grafting model beliefs across checkpoints
Excerpt
arXiv:2610.00767v1 Announce Type: new Abstract: Pre-training interventions are critical to alignment research, since beliefs formed during pre-training shape how a model generalizes from later training. One recently popular technique for such interventions is synthetic document fine-tuning (SDF), which aims to alter what the model believes. Ideally, synthetic documents would be mixed into pre- or mid-training, but every change to a pre-training corpus must be followed by a full post-training run