← Back to all articles
Reddit r/MachineLearningSeptember 22, 2026

Simulating fault tolerance with stage skipping in pipeline-parallel training [R]

Excerpt

Our most recent work at Templar explores fault tolerance in Crucible, our distributed pre-training platform. The goal is to keep healthy workers training when another pipeline stage goes offline. Crucible combines data-parallel replicas with pipeline parallelism. Each replica holds a copy of the model, split into stages on separate workers. SparseLoCo exchanges compressed updates between replicas, while pipeline compression reduces the communication across stage boundaries. We combine those meth