← Back to all articles
arXiv cs.CLSeptember 28, 2026

UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification

Excerpt

arXiv:2605.06221v2 Announce Type: replace Abstract: As large language models (LLMs) continue to advance rapidly, they are becoming increasingly capable while simultaneously demanding ever-longer context lengths. To improve the inference efficiency of long-context processing, several novel low-complexity hybrid architectures have recently been proposed, effectively alleviating the computational burden of long-context inference. However, existing research on long-context prefill acceleration remai