← Back to all articles
arXiv cs.LGOctober 7, 2026

AEGIS: Runtime-Guided GPU Collocation for Multi-Tenant Deep Learning Training

Excerpt

arXiv:2508.19073v4 Announce Type: replace-cross Abstract: Deep learning training commonly runs on shared multi-tenant GPU servers, where exclusive allocation provides isolation but can leave resources underutilized and increase queueing time. Collocation can improve efficiency, but interference-agnostic placement may cause severe slowdowns, while inaccurate memory information can lead to out-of-memory (OOM) failures. We present AEGIS, a server-scale runtime scheduling system for controlled collo