Reddit r/LocalLLaMASeptember 20, 2026
Speed-up Kimi K3(2.8T) on a 16x GB10 Cluster — 30 t/s coding throughput, 136 t/s concurrency peak.
Excerpt
I wanted to share a quick update and performance video running the full Moonshot AI Kimi K3 (moonshotai/Kimi-K3) model across my 16x GB10 cluster. Getting a 2.8T parameter model running smoothly requires custom runtime patches and a solid network layout, but throughput and concurrency on coding tasks have been impressive. Performance Benchmarks Coding Generation / Decode: Sustaining ~30 tok/s (peaking around 38 tok/s) during heavy code generation and agentic tasks. Prefill Throughput: ~75