Reddit r/LocalLLaMAAugust 19, 2026
I pushed Qwen3.8-27B limits again... Dflash2 - 134 tps on a RTX 3090
Excerpt
Edit: Title says 134 tps, it's actually 138 -- keep in mind my 3090 is power limited to 250w. Three days ago I released a hyper-optimized Qwen3.8-27B inference engine for an RTX 3090 (82 tps single request, 672 peak), and yesterday's update took it to ~114 tps single-user / ~1,000 tps at 64 concurrent. Today it's ~138 tps at default sampling on real chat prompts (up from ~124), 942 tps at 64 concurrent (re-measured today on the current stack), and the thing I'm actually happy about: a follow-up