← Back to all articles
Reddit r/LocalLLaMASeptember 9, 2026

Solved: LLM inference on Windows was 2–3x slower when the server window wasn't focused

Excerpt

The fix: run the server detached/headless instead of keeping it attached to a console window. RTX 5090, ~27B NVFP4 model via ninfer: Terminal focused: 130–200 tok/s Terminal unfocused: 50–60 tok/s Click the terminal → immediately back to 130+ tok/s At first I thought GPU throttling, but the GPU wasn't the problem. During the slow runs: SM clock was actually higher : 2550–2600 MHz vs 2100–2200 MHz Power limit was the same: 400 W No PCIe power-state drop decode-host stayed around 390–494 µs The bi