← Back to all articles
Reddit r/LocalLLaMAAugust 31, 2026

Qwen3.8-Flash-Next in llama.cpp from CPU-only to 96GB VRAM: 8.5 to 109 tok/s, max context and parameters test. My findings on RTX 6000 PRO.

Excerpt

Hey guys, I tested Qwen3.8 Flash with llama.cpp from CPU-only to the full 96GB of my RTX PRO 6000. Short version: CPU-only reached 8.34 tok/s at a 2K prompt Full 96GB reached 109.07 tok/s At 245K context, 24GB to 96GB gave 14.89 to 21.61 tok/s The 96GB advantage over 24GB decreased from 2.80x at 2K to 1.45x at 245K Forcing the 27.2 GiB PLE table onto CUDA reduced decode from 108.5 to 1.95 tok/s RAM-resident loading gave 1.87x more prefill than mmap Non-unified KV reached 92.0 tok/s total output