Reddit r/LocalLLaMAAugust 26, 2026
Qwen3.8 27B C8 at 972 TG / 5,680 PP on 4x MI100 rig ($6.5k) using my new INT8 vLLM fork
Excerpt
Yet another vLLM fork thread here, but this time its for older INT8-centric hardware. This is a complete INT8 serving stack for Qwen3.8 27B based on vLLM, AITER, and a 27B GPTQ INT8 quant w/ DFlash2 . Its not just another vibed autoresearch loop. No, vLLM ships with very little int8 support, and this stack adds INT8 into every crevice of Qwen3.8 including in dependent libraries and new fused kernels. So no longer are your old INT8-centric cards relegated to second rate algos and suboptimal dtype