arXiv cs.AIOctober 7, 2026
Characterizing Parallelism Strategies in LLM Inference: Fundamental Compute-Communication Trade-offs
Excerpt
arXiv:2610.05305v1 Announce Type: cross Abstract: Large Language Model (LLM) inference has become the dominant workload in modern AI systems, requiring serving infrastructures to maximize throughput while meeting strict latency Service-Level Objectives (SLOs). Since state-of-the-art LLMs exceed the compute and memory capacity of a single GPU, inference is commonly distributed across multiple GPUs using tensor parallelism (TP), pipeline parallelism (PP), or hybrid parallelism (HB). However, selec