← Back to all articles
Reddit r/LocalLLaMASeptember 20, 2026

The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks

Excerpt

TL;DR: Local agent loop, ~21 days, one RTX 3090. Task was pretty much "build a CUDA inference engine for optimized for yourself on this GPU arch." Got working kernels and benches, not a win over llama.cpp. ~12 human messages. Compaction ate ~83 hours. Old joke: you don’t criticize how well the bear dances, you’re surprised it dances at all. Setup: Qwen 3.8 27B Q4, Q8 KV, 200k context, deepseek harness, written rulebook: roles, handoffs, when to ping me, don't copy llama.cpp, don't declare the ta