Reddit r/LocalLLaMAAugust 20, 2026
3 days benchmarking most llama.cpp flags on my weird 40gb vram laptop + tb4 egpu setup. Got +70% generation, +40% prefill, 60k more context, and filed a bug in llama around MTP. What I learned.
Excerpt
tldr: went from 16~ t/s to 27~ t/s generation. got my usable context up from 220k to the full 262k without sacrificing anything. prefill also increased from 376 to 573 command I ended up with, fwiw: llama-server -m Qwen3.8-27B-UD-Q6_K_XL.gguf -c 262144 -ngl 999 -fa on \ -ctk q8_0 -ctv q8_0 --spec-type draft-mtp,ngram-map-k4v --spec-draft-n-max 3 \ --spec-draft-device CUDA0 --split-mode layer -dev CUDA0,Vulkan2 -ts 40,60 \ --jinja -fitt 256 Note that my setup is a bit atypical. P1 Gen 6 (an RTX 4