Reddit r/LocalLLaMASeptember 11, 2026
Hot Expert Reload on GPU is what this community needs
Excerpt
A huge favor to ask llama maintainers - please implement this feature. Even with one 3090 card there will be tangible improvements in decode speed on MOE models with moderate number of active parameters, like Qwen3.8-Flash-Next, Deepseek V4/V4.1 Flash, GLM 5.3 Flash. With 2x 3090 speeds will be quite close to the full offload of these models to VRAM. This will make these almost SOTA models really usable locally. submitted by /u/perelmanych [link] [comments]