Reddit r/MachineLearningSeptember 22, 2026
Simulating fault tolerance with stage skipping in pipeline-parallel training [R]
Excerpt
Our most recent work at Templar explores fault tolerance in Crucible, our distributed pre-training platform. The goal is to keep healthy workers training when another pipeline stage goes offline. Crucible combines data-parallel replicas with pipeline parallelism. Each replica holds a copy of the model, split into stages on separate workers. SparseLoCo exchanges compressed updates between replicas, while pipeline compression reduces the communication across stage boundaries. We combine those meth