Reddit r/LocalLLaMASeptember 22, 2026
Qwen 3.8 27B at ~3 BPW on an RTX 3060: GSQ vs ByteShape IQ3-XXS 2.88BPW
Excerpt
Someone recommended that I try the ByteShape Qwen 3.8 27B IQ3-XXS GGUF after seeing my previous testing of the GSQ quant. So I did. And the result was… surprisingly bad. For context, I'm running: RTX 3060 12GB 16GB DDR4 RAM, single channel CachyOS / Arch Linux llama.cpp Qwen 3.8 27B MTP/speculative decoding where applicable The two low-bit quants I compared were: ISTA-DASLab / GSQ-RCO-IQ3-XXS ~10.4GB roughly 2.5 BPW territory MTP enabled ~29 tok/s around full context ~34–40 tok/s at lower contex