← Back to all articles
Reddit r/LocalLLaMAAugust 22, 2026

GLM-5.2 local inference: ubatch size made a much bigger difference than I expected

Excerpt

Has anyone else experimented with ubatch size when running GLM-5.2 locally? I was testing the 226 GiB GLM-5.2-UD-IQ2\_XXS GGUF on 3x RTX PRO 6000 Blackwell GPUs and got a pretty interesting result. With llama.cpp vs my TensorSharp implementation: llama.cpp TS ubatch 1024 TS ubatch 2048 pp128 276.5 254.8 264.4 pp512 695.4 666.9 659.6 pp2048 763.1 918.9 1145.8 pp4096 715.8 864.7 1048.7 tg64 42.2 43.7 43.9 All numbers are tokens/sec and were measured back-to-back on the same machine. What surprised