← Back to all articles
Reddit r/LocalLLaMASeptember 23, 2026

DeepSeek-V4-Flash-0731 at ~40–50 tok/s on 2× Radeon AI PRO R9700 with the affinity engine (prebuilt quant + fixes)

Excerpt

I've been running DeepSeek-V4-Flash-0731 on two Radeon AI PRO R9700s (32 GB each, 192 GB system RAM) using affinity( https://codeberg.org/StillDeadcode/affinity ), an inference engine written by **Yoshi Exeler (StillDeadcode)** specifically for DeepSeek-V4-Flash on one or two RDNA4 cards. All the credit for the engine goes to them. It keeps the hot experts on the GPUs, streams the rest from host RAM, and runs a DeepSeek-native speculative draft. If you have R9700s, go look at it. What I added: -