Reddit r/LocalLLaMASeptember 19, 2026
Improved TPS of Gemma 4 31B : the journey and also creating custom patches with VLLM fork
Excerpt
I improved the TPS of Gemma 4 31B. Improving TPS and performing optimisations requires understanding of the model architecture, and I had to fork VLLM and apply my own patch to break into making a configuration work as per my idea. I wrote a full article so that even beginners can understand and improve TPS for any model, this will serve as a mental model for anyone who start with TPS optimisations. Also I strongly believe being just a mecha pilot wont be enough for inference optimisations. I ha