← Back to all articles
Reddit r/LocalLLaMAAugust 22, 2026

Tesla P40 - use F16 KV instead of Q8

Excerpt

If you use these cards together, you would think Q8 would be faster tps because it uses less VRAM. Well the reality is: Prompt TPS is nearly identical between the two (e.g. Q4@45k: 237.7 vs 237.6), so KV type doesn't affect prefill. The divergence is purely in generation, where every token requires reading the full KV cache for attention across all 40+ layers. The cost of q8_0: - Per-token dequantization: Every attention read must convert q8_0 → f16 on-the-fly before the matmul. That's ctx × n_h