Reddit r/LocalLLaMAAugust 26, 2026
OpenCode with Qwen3.8-27B for Small Games or Browsing the Web With 16GB VRAM
Excerpt
In the past, I have use llama.cpp, but I read that the exl3 quantization format should give better precision , so I have tried exllamav3/tabbyAPI. It was able to write the shown simple HTML game without interaction after asking some questions. The following was tested on a laptop with a NVIDIA RTX A5000 laptop (16 GB) GPU. With the 3 bpw model and 6 bit/5 bit KV cache, the maximum context length is around 110k tokens with MTP. This gives around 55 tokens/s decode speed for code and around 10 tok