← Back to all articles
arXiv cs.LGOctober 2, 2026

Denoising Surface: Modeling and Predicting Inference Cost for Diffusion LLM Serving

Excerpt

arXiv:2610.00499v1 Announce Type: new Abstract: As diffusion large language models (dLLMs) become more capable, they are moving from research settings to real-world \textit{serving}, where request management (such as scheduling and resource allocation) relies on accurate estimation of per-request inference cost. However, common cost proxies fall short for dLLMs: output length ignores that one forward pass can unmask multiple tokens, and denoising-step count ignores the \textit{heterogeneous} per