← Back to all articles
arXiv cs.CLSeptember 21, 2026

VQ-Logits: Compressing the Output Bottleneck of Large Language Models via Vector Quantized Logits

Excerpt

arXiv:2505.10202v2 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved remarkable success but face significant computational and memory challenges, particularly due to their extensive output vocabularies. The final linear projection layer, mapping hidden states to vocabulary-sized logits, often constitutes a substantial portion of the model's parameters and computational cost during inference. Existing methods like adaptive softmax or hierarchical softmax introduce struct