arXiv cs.CLSeptember 11, 2026
Why Does Post-Training Quantization Work?
Excerpt
arXiv:2609.11716v1 Announce Type: cross Abstract: Post-training quantization compresses large language models (LLMs) by storing their weights at reduced precision, and each quantized weight introduces an error into the hidden states. Naively, these errors should accumulate with depth and corrupt next-token prediction; randomly initialized models accumulate these discrepancies rapidly, whereas quantized pretrained models accumulate much less hidden-state error and largely maintain downstream task