← Back to all articles
Reddit r/LocalLLaMASeptember 22, 2026

People keep asking how I'm running Qwen3.8 27b @ 160k context on a 16 GB card, so here's the whole thing. No tricks, no 8-bit KV hacks, no CPU offload party trickery — just a quirk of the model's architecture and a couple of flag choices.

Excerpt

I asked Qwen3.8 how this is done, because it works so well honestly it's like having a local claude. I have done a few shot prompts, like a multiplane flight sim etc. dam amazing even at this low quant. so here are my details per Qwen3.8 27b on how it's setup. so meta. lol FYI I use PI coder ## The rig - **GPU:** AMD Radeon RX 7600 XT, 16 GB (Navi 33 / gfx1102) - **CPU:** i7-8700 (6c/12t), 32 GB DDR4 - **Driver path:** Vulkan (RADV), not ROCm — `--device Vulkan1` because `Vulkan0` is my Intel iG