Reddit r/LocalLLaMASeptember 22, 2026
Qwen3.8-27B: >70 tok/s (>160 tok/s concurrent), 10k tok/s prefill, full context on 2x3090 (or and 48GB or larger on ampere or higher), vanilla vllm
Excerpt
I didn't know my set up was outperforming nearly everyone until reading another discussion where people were struggling getting half of that speed with half the context on the same hardware. I benchmarked a couple dozen quants, vllm, sglang, llama.cp and benchmarked settings and configurations on each to find the following: I can get full 262k context (can nearly fit 2 full contexts in kv cache simultaneously), >70 tok/sec single stream decode, close to 200 tok/sec decode at 3 to 4 concurrency (