Reddit r/LocalLLaMAAugust 27, 2026
Qwen3.8-Flash-Next: Time to Update Those Benchmarks
Excerpt
specs hardware: M4 Max 128GB Studio inference engine: oMLX & lllama.cpp insights it still very early, so had to disable oMLX K/V caching, qwen4_exp architectureis not yet supported + the obvious n-grams with which the whole 4 bit quant takes ~100G, so pretty tight nevertheless, this is the first model for the year that was able to break through 94% on my cupel benchmark one interesting bit is Qwen 3.8 27B is obviously great, but it did not do that well, since I have coding, general knowledge and