Reddit r/LocalLLaMASeptember 22, 2026
MiMo-V2.6-Flash on vLLM: fixes for "empty responses" with thinking + tools, and a hidden 2,048-token output cap
Excerpt
Some people here say MiMo-V2.6 is bad with tools and are going back to GLM-5.3-Flash. I spent today running MiMo-V2.6-Flash-RL as the backend for an agent harness, on 2× DGX Spark with vLLM, using the tonyd2wild recipe. Most of the "tool problems" I hit turned out to be serving bugs rather than the model. There are three separate issues, all fixable without touching vLLM's core logic. Details below, in case it saves someone a day. 1. With thinking on, replies after the first turn come back "empt