← Back to all articles
arXiv cs.LGOctober 7, 2026

Activation Denoising: A Robustness View on Parallel vs Sequential LLM Quantization

Excerpt

arXiv:2610.07522v1 Announce Type: new Abstract: Post-training quantization is a powerful tool for compressing large language models. The most scalable methods quantize every layer in parallel, but quantization errors then compound through the residual stream, as no layer corrects for the errors of the layers before it. Sequential quantization accounts for this error compounding by re-calibrating each layer on the already-quantized outputs of its predecessors, yielding stronger results but at the