← Back to all articles
arXiv cs.AIOctober 2, 2026

HHR: Hierarchical Hash Retrieval for Efficient LLM Generation

Excerpt

arXiv:2610.01230v1 Announce Type: new Abstract: Efficient long-context inference is essential for large language models (LLMs), yet it poses a severe computational bottleneck. Hash-based retrieval offers an efficient alternative by encoding queries and keys into binary codes and using Hamming distance for key selection. However, this leads to a critical mismatch between Hamming distance and attention relevance. Query-Key logits depend jointly on directional similarity and feature magnitudes, whe