← Back to all articles
arXiv cs.LGOctober 7, 2026

SoloQ: Calibration-Free Quantization for Diffusion Language Models

Excerpt

arXiv:2610.07121v1 Announce Type: new Abstract: Diffusion large language models dLLMs) have emerged as a promising alternative to autoregressive language models through bidirectional diffusion-based token generation. However, their growing model sizes and high inference costs make efficient deployment challenging: full-sequence denoising repeatedly invokes compute-intensive forward passes, while block-diffusion models additionally introduce a memory-intensive KV-cache. Low-bit weight-activation