← Back to all articles
arXiv cs.CLSeptember 14, 2026

SeDeM: Selective Decompression of Hidden-State Memories for Long-Context Question Answering

Excerpt

arXiv:2608.00311v2 Announce Type: replace Abstract: Long-context inference with large language models (LLMs) is costly: self-attention during prefill scales quadratically with sequence length, and the key-value (KV) cache grows with the number of processed tokens. Larger context windows also do not ensure reliable evidence use. Context compression reduces this cost, but many soft-compression methods use LLMs as compressors and rely on compact memory tokens both to preserve information and to con