← Back to all articles
Reddit r/LocalLLaMASeptember 2, 2026

First local-LLM tuning attempt: Qwen3.8-27B true Q4_K_M at 13.2 tok/s near 50-61K context on RTX 5080 16GB

Excerpt

This was my first serious attempt at tuning a local LLM. I started because Qwen3.8-27B IQ3 was fast on my RTX 5080 but the coding quality disappointed me, and the Q4 profiles I tried in LM Studio were much slower than reports here. Hardware: RTX 5080 16 GB i5-14600K 64 GB DDR5-5600 (4 DIMMs) Windows Final model/runtime: Unsloth Qwen3.8-27B UD-Q4_K_M, unmodified (16.46 GB) official llama.cpp b10760 CUDA 13.3 build 65,536 context, one slot Q4_0 K/V cache, Flash Attention medium thinking, text only