← Back to all articles
Reddit r/LocalLLaMASeptember 11, 2026

llama.cpp ngram on RAM/SSD?

Excerpt

I've been out of the loop for some time. Is there already an official way to offload ngram to RAM or SSD in something like Unsloth Studio? Interested in running Qwen3.8-Flash-Next on 72GB VRAM, but naiive attempts failed because even at Q4 it seems to load God knows what to God knows where. Would appreciate if someone share the recipe, or tell what are the official plans to support this (I can wait, knowing that it is upcoming). Thank you in advance. submitted by /u/NickNau [link] [comments]