← Back to all articles
Reddit r/LocalLLaMASeptember 9, 2026

What settings do you use for running Qwen3.8-Flash-Next in llama.cpp?

Excerpt

Hi, I'm wondering what settings you are using in order to run Qwen3.8-Flash-Next on your devices? I'm especially interested in setups with 96GB VRAM. I'm not quite sure if llama.cpp does offload the embeddings to RAM or disk with my settings. I would like to offload them to RAM in order to avoid too much performance penalty. These are the settings I use and which work the best at the moment: [qwen3.8-flash-next] model = /mnt/kyouma/1TB/ML/models/unsloth/Qwen3.8-Flash-Next-GGUF/UD-Q4_K_XL/Qwen3.8