โ† Back to all articles
Reddit r/LocalLLaMASeptember 11, 2026

CUDA/HIP: Flash Attention tuning (gfx1201) by pwilkin ยท Pull Request #28102 ยท ggml-org/llama.cpp

Excerpt

Nice pp improvements for RDNA4(R9700) & 3.5(RX 9060 XT, 8060S). More good numbers on large context. PR has detailed benchmarks. u/ilintar ๐Ÿ‘ submitted by /u/pmttyji [link] [comments]