arXiv cs.LGOctober 7, 2026
AEGIS: Runtime-Guided GPU Collocation for Multi-Tenant Deep Learning Training
Excerpt
arXiv:2508.19073v4 Announce Type: replace-cross Abstract: Deep learning training commonly runs on shared multi-tenant GPU servers, where exclusive allocation provides isolation but can leave resources underutilized and increase queueing time. Collocation can improve efficiency, but interference-agnostic placement may cause severe slowdowns, while inaccurate memory information can lead to out-of-memory (OOM) failures. We present AEGIS, a server-scale runtime scheduling system for controlled collo