← Back to all articles
Reddit r/LocalLLaMASeptember 3, 2026

Qwen-3.8-Next-Flash Ngram Hot-Swappable Knowledge Injector for llama.cpp

Excerpt

Looking into the new Qwen architecture, I was curious if you could modify the Ngram PLE Table to make it work like a long-term knowledge database. It turns out that, with some limitations, you can. I coded a small modification to llama.cpp to modify the table in-memory, allowing you to patch it with new data in real time. The PLE table is updated on every prompt, so now you can hot-swap parts of it without reloading the model. The limitation is that it’s hard to control the output reliably, as t