Reddit r/LocalLLaMASeptember 9, 2026
GLM 5.3 Flash Q4 @ 60tps / 550tps on M3 Ultra
Excerpt
I have been a dwarfstar fan for awhile and I really liked glm 5.3 flash but needed it to be materially faster to feel good using it. In the screenshot you can see the outcome of using the model with a claude code harness at ~200k depth, with many tool calls and averaging over 38tps output. Yes, I put 60tps in the headline and you will get that if you ask it to write SQL. https://github.com/IngeniousIdiocy/ds4/blob/glm53-m3ultra/README.md#glm53\_m3ultra Main ds4 was single-stream serially decodin