arXiv cs.LGOctober 1, 2026
Efficient Expert-Parallel Communication on PCIe-Connected Consumer GPUs
Excerpt
arXiv:2609.40093v1 Announce Type: cross Abstract: Expert parallelism (EP) enables inference of large Mixture-of-Experts (MoE) models by placing their experts across multiple GPUs, but requires substantial communication between GPUs at every MoE layer. As contemporary MoE models activate more experts per token, this communication accounts for a growing fraction of inference time. The cost becomes particularly pronounced on PCIe-based consumer GPU systems, where all inter-GPU transfers traverse CP