arXiv cs.LGAugust 18, 2026
Beyond Binary Priorities: Multi-Tier SLA Scheduling for Large Language Model Serving
Excerpt
arXiv:2608.16336v1 Announce Type: cross Abstract: Modern LLM serving deployments must simultaneously satisfy heterogeneous service-level objectives (SLOs) across a diverse population of user tiers, ranging from latency-critical API calls to background batch processing. Llumnix introduced a dynamic, migration-capable multi-instance scheduler for LLM inference that achieves load balancing, defragmentation, prioritization, and auto-scaling through a unified "freeness" metric. However, Llumnix's pri