← Back to all articles
Reddit r/LocalLLaMASeptember 18, 2026

Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides

Excerpt

I had Astra run a multi-hour investigation into speeding up local MoE inference while keeping the model weights and quantization unchanged. The experiments covered Qwen3.6-35B-A3B and Qwen3.8 Flash-Next. Some workloads showed substantial gains, particularly prompt processing and source-code editing. There were also regressions and tradeoffs, which are documented alongside the results. Hardware: RTX 4080, 16 GB VRAM Ryzen 9 5900X 64 GB DDR4 RAM I’ve shared the source patches, benchmark summaries,