Reddit r/LocalLLaMASeptember 13, 2026
Qwen 3.8 27B UD-IQ4_XS even faster on 16GB CUDA
Excerpt
This is an evolution on top of Raymond's KV cache streaming fork - all credits to what enabled this goes to him. The basic idea behind what he enabled was a pool of memory in VRAM that is used differently depending on the phase (prompt processing or decoding) and when total used context is larger than what fits in VRAM it's instead streamed from host RAM in time for when the current layer needs it. It enables _much_ higher TG tps than regular llama.cpp offloading to host RAM. I've used it to run