← Back to all articles
Reddit r/LocalLLaMAAugust 20, 2026

[Draft - Open PR] AVX2: Speed up large batch size prompt processing of IQ models by bartowski1182 · Pull Request #27402 · ggml-org/llama.cpp

Excerpt

IQ quants are particularly slow on CPU at large batch sizes (what you'd see for imatrix and perplexity) Benchmark numbers I ran PPL against master and this PR to get speed and numbers on --chunks 50 for Qwen3.6-27B and Qwen3.6-35B-A3B on EPYC 9654 using 24 threads Created pure IQ1_S , IQ1_M , IQ2_XXS , IQ2_XS , IQ2_S , IQ3_XXS , IQ3_S , IQ4_XS , and IQ4_NL . Made pure to make sure each tensor type is fully exercised. These are the most extremely differences because it's at a big batch size (512)