arXiv cs.AIOctober 7, 2026
Greedy Local Learning for Language Model Pretraining: Gaps and Objective Design
Excerpt
arXiv:2610.04867v1 Announce Type: cross Abstract: Greedy block-wise local learning splits a network into gradient-isolated blocks trained by local auxiliary losses, deleting the backward pass between blocks: inter-stage communication becomes forward-only and every block can step its optimizer independently, properties directly relevant to decentralized model-parallel training. Local learning is competitive with end-to-end backpropagation on image classification, and on small Transformers it is k