Reddit r/LocalLLaMASeptember 7, 2026
Has anyone noticed a difference between bf16 and q8 quantization ever?
Excerpt
I'm currently running qwen 3.8 27b q8 and super happy with it. I have a DGX spark so could theoretically run at bf16 quantization. I know the measurable differences between an 8 bit and 16 bit quant are small, but I guess I get fomo, like 1% of tokens differ, but what if those are the hardest most important tokens? I guess I'm just getting fomo over bs16 and wondering if other people have tried it and noticed a difference? submitted by /u/superSmitty9999 [link] [comments]