← Back to all articles
Reddit r/LocalLLaMASeptember 27, 2026

... so, yeah.

Excerpt

Finally got 3.8-Flash-Next running on my M4Pro 48GB Mac with https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF Dense 3.8-27B is just faster... and maybe better due to quantization level... EDIT: Hold a second, Flash-Next is actually performing faster than 27B after some key flags on llama.cpp. it's Holding up to 131K without OOM-ing..... maybe... 0.36.940.283 I srv load: --top-k 0.36.940.283 I srv load: 20 0.36.940.284 I srv load: --ctx-size 0.36.940.284 I srv load: 131072 0.36.