← Back to all articles
Reddit r/LocalLLaMAAugust 19, 2026

Qwen3.8-27B on 2x 3090 + vLLM + DFlash2: 218 tok/s single request

Excerpt

I hacked this together so there's probably more on the table in terms of performance. Measured with the Club-3090 canonical bench suite (bench.sh, 3 warmups + 5 measured runs, temp 0.6 / top_p 0.95 / top_k 20). Prefill: 1342 tok/s @ 10k, 628 tok/s @ 90k Spec-decode: 7 draft tokens, acceptance length 3.35, 47.8% acceptance Peak VRAM: 22.3 GB/card Context ceiling: 131k (DFlash2 drafter eats ~13.5 GB) Used Kimi K3 for all the VLLM fixes Metric Narrative Code Decode TPS 120.1 218.3 Wall TPS 117.7 20