Reddit r/LocalLLaMASeptember 4, 2026
Qwen3.8-Flash-Next: 256k context, 16tok/s on DDR4 and a Tesla T4
Excerpt
I've got an refurb Dell R740 running Proxmox that I put a Tesla T4 in, mainly to run some CTC local transcription work, but thought it would be fun to try DS4 when it came out, and it was appalling at around 2 tok/s. However pulled it out again when Qwen3.8 dropped, and it was much improved, particularly with ik_llama. Hardware: Dell R740, 2x Xeon Gold 6230, 384GB DDR4-2666, one Tesla T4 16GB. Guest VM pinned to one NUMA node: 20 cores, 168GB RAM. Model: Unsloth Qwen3.8-Flash-Next UD-Q4_K_XL, 11