← Back to all articles
Reddit r/LocalLLaMAAugust 30, 2026

Qwen3.8-Flash-Next turns 4xR9700 into a local AI powerhouse! 120 t/s TG and 12k t/s PP single request with optimized vLLM

Excerpt

If you own 4xR9700 and were waiting for the model to make them shine, then I have some good news for you! It's running at 80-120 tokens/second for generation and 12k token/second prefill for a single request, using tcclaviger's MXFP4-FP8 quant and custom vLLM image docker.io/tcclaviger/vllm:DevQwenNextFlash optimized for R9700. Total context (shared across all parallel requests) in this setup is 700k tokens. Here is the full command: podman run --rm -it \ --init \ --network host \ -v /models:/mo