Reddit r/LocalLLaMASeptember 12, 2026
Qwen3.8 Flash Next now at 1.2k t/s prefill on Strix Halo
Excerpt
As you all know, Qwen3.8 Flash Next on mainline llama.cpp is still in a pretty experimental stage, but a lot of community forks are trying to get it to work better. There's also a closed-source solution called Halogen ( https://github.com/peonist-ai/halogen-flash-server ) that boasted 1.2k t/s prefill numbers when the community fork barely reached 400. Since I dislike closed source and I like open source, I decided to take the challenge and bring llama.cpp up to the same performance level and I'