← Back to all articles
arXiv cs.AIOctober 2, 2026

TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference

Excerpt

arXiv:2610.01763v1 Announce Type: new Abstract: Activation sparsity speeds up large language model (LLM) inference by setting unimportant activations to zero so that the corresponding computations can be skipped. Existing training-free methods, however, make different trade-offs: threshold-based methods such as TEAL adapt the sparsity level to each token but do not tightly control the realised sparsity, while TopK-based methods such as WINA enforce a fixed sparsity level but use the same sparsit