Reddit r/LocalLLaMAAugust 20, 2026
Qwen 3.8 27B KV f16 vs q8_0 are not equivalents
Excerpt
I'm testing it since release, now with UD 3.0 in my AMD R9700 with ROCm, I always read everywhere that F16 and q8_0 for KV cache are essentially the same... well, I tested it and I can see differences. Some differences are minimal, F16 is more careful and detailed with structured and free output, thinking process almost all the time on point, and after 120k ctx can keep delivering as it was at I saw some yt videos and posts with bad/mixed reviews, reading the details, q4_0 KV cache... ouch Anyon