← Back to all articles
Reddit r/LocalLLaMASeptember 12, 2026

DeepSeek V4.1 Flash on 8× A40: ~40 tok/s Q2_K and ~32 tok/s Q4_K_M with TensorSharp

Excerpt

Hey everyone — I’m building TensorSharp , an open-source LLM inference engine. Here are the latest DeepSeek V4.1 Flash GGUF results using its native ggml_cuda backend. Setup: 8× NVIDIA A40, layer split, F16 KV cache, 65,536-token configured context. Prefill measurements use approximately 4.9K-token prompts—not the full context window. Final optimized results — all speeds in tokens/sec: Metric | Q2_K | Q4_K_M Prefill | 533–539 | 451.8–492.1 Single-request decode | 40.31–40.72 | 31.0–32.5 Decode,