arXiv cs.LGOctober 1, 2026
LampAttention: Look-Ahead Mixed-Precision FlashAttention for Dedicated Accelerators
Excerpt
arXiv:2609.39361v1 Announce Type: new Abstract: While most attention logits can be computed in low precision without degrading numerical stability, current attention kernels fail to exploit this phenomenon. We introduce a novel hardware-algorithm co-design in the form of mixed-precision FlashAttention. Our method accumulates key-query products and evaluates their exponentials in 8-bit formats, then adaptively identifies sensitive sub-blocks and recomputes them in 16-bit formats. We propose the s