Reddit r/LocalLLaMASeptember 11, 2026
Qwen3.8 Flash Next UD-Q4_K_XL 49 tokens/s TGS using 2x RTX 3090 on Windows 11.
Excerpt
Developer's own thread: https://www.reddit.com/r/LocalLLaMA/s/adp1cGZZe9 Code: https://github.com/Inovello/llama.cpp/tree/flashnext-e06 My hardware: 2x RTX 3090, Intel Ultra 7 270k Plus, 192 GB DDR5@5600 MHz Token generation speed went from 20 t/s to 49 t/s. Prompt processing speed is 140 t/s. Prompt processing is faster on the main branch. https://preview.redd.it/an5rtqz5nwoh1.png?width=643&format=png&auto=webp&s=6b9f8760d2440e4a40420f27956179808826017a I have CUDA 13.3.1 installed. I use Windo