Reddit r/LocalLLaMAAugust 22, 2026
GLM-5.2 local inference: ubatch size made a much bigger difference than I expected
Excerpt
Has anyone else experimented with ubatch size when running GLM-5.2 locally? I was testing the 226 GiB GLM-5.2-UD-IQ2\_XXS GGUF on 3x RTX PRO 6000 Blackwell GPUs and got a pretty interesting result. With llama.cpp vs my TensorSharp implementation: llama.cpp TS ubatch 1024 TS ubatch 2048 pp128 276.5 254.8 264.4 pp512 695.4 666.9 659.6 pp2048 763.1 918.9 1145.8 pp4096 715.8 864.7 1048.7 tg64 42.2 43.7 43.9 All numbers are tokens/sec and were measured back-to-back on the same machine. What surprised