arXiv cs.LGOctober 1, 2026
SparseEngine: Sparse-First Inference Engine
Excerpt
arXiv:2609.39068v1 Announce Type: new Abstract: Long-context LLM agents accumulate interaction histories that strain KV-cache memory and attention computation. Although sparse attention reduces these costs, heterogeneous cache representations and workflows hinder integration with existing inference engines, while prior sparse-serving abstractions support only specific layouts or workflows. We present SparseEngine, a ground-up, sparse-first inference engine whose shared lifecycle contract lets ea