arXiv cs.LGOctober 1, 2026
Characterizing High Bandwidth Flash for LLM Serving
Excerpt
arXiv:2609.39131v1 Announce Type: new Abstract: Large language model (LLM) serving requires substantial memory to store model weights and KV caches. As models grow larger and contexts become longer, memory capacity and bandwidth increasingly become bottlenecks for serving performance. Agentic workloads compound this pressure through repeated interactions over growing contexts, making it increasingly important to retain KV state for reuse. High-bandwidth flash (HBF) offers a way to expand acceler