← Back to all articles
arXiv cs.LGOctober 1, 2026

RATIO: Reasoning Analysis and Token-level Inference Optimization for Quantized Reasoning Models

Excerpt

arXiv:2609.39801v1 Announce Type: new Abstract: Post-training quantization (PTQ) has become a widely adopted technique for reducing the memory footprint and inference cost of large language models (LLMs). However, recent studies reveal that when applied to reasoning models, PTQ not only degrades reasoning performance but also exacerbates overthinking, leading to longer reasoning trajectories. These issues may offset the efficiency gains expected from lower-precision inference. Existing approache