← Back to all articles
Reddit r/LocalLLaMASeptember 19, 2026

Qwen-3.8-Flash-Next on 1x RTX 5090: TG=50 t/s, PP=2300 t/s - with FreeToken

Excerpt

I was somewhat disappointed with the performance of Qwen 3.8 Flash Next on my single RTX 5090 using llama.cpp. One issue is that llama.cpp still has no Expert Caching implemented for MoE models. There are various PRs and discussions (see here , for example), but nothing is merged yet. Then I stumbled upon FreeToken which claims to "unlock datacenter-class intelligence on the hardware you already own". With that I am able to run 3.8 Flash Next with TG ~ 50 t/s (around 40-60 depending on how well