arXiv cs.LGOctober 1, 2026
MatGPTQ: Efficient and Accurate Inference over Nested Quantized Models
Excerpt
arXiv:2602.03537v2 Announce Type: replace Abstract: Matryoshka Quantization (MatQuant), Any-Precision-LLM (AP) and AnyBCQ (AB) are recent quantization approaches showing that a single integer-quantized model can be served across multiple precisions. In this paradigm, lower-precision models are extracted from a higher-precision model by simply reading fewer bits of the weights. This enables a single checkpoint to cover a wide range of memory and latency budgets, but makes both quantization and ef