arXiv cs.AIOctober 7, 2026
Dissecting Quantization Error: A Concentration-Alignment Perspective
Excerpt
arXiv:2603.04359v2 Announce Type: replace-cross Abstract: Quantization can drastically increase the efficiency of large language and vision models, but typically incurs an accuracy drop. Recently, function-preserving transforms (e.g. rotations, Hadamard transform, channel-wise scaling) have been successfully applied to reduce post-training quantization error, yet a principled explanation remains elusive. We analyze linear-layer quantization via the signal-to-quantization-noise ratio (SQNR), show