← Back to all articles
arXiv cs.AIOctober 2, 2026

The Devil Is in the Reconstruction Loss Scale: Rethinking Optimization in LLM Quantization

Excerpt

arXiv:2610.00983v1 Announce Type: cross Abstract: Post-training quantization (PTQ) methods typically use sequential quantization that partitions a pre-trained LLM into a series of units (e.g., transformer blocks), with one unit quantized at each stage. State-of-the-art PTQ methods are predominantly learning-based, optimizing auxiliary quantization parameters (e.g., scaling factors, rotation matrices, clipping thresholds, and adapters) via gradient descent to minimize a reconstruction loss. A com