← Back to all articles
Reddit r/LocalLLaMASeptember 3, 2026

Introducing Quartermaster, an open source local AI platform designed for ease of use that does not sacrifice customizability

Excerpt

It started as a fork of llama-swap , but I have been building it out for myself since then as a convenient tool for all my local AI needs, and by now it has drifted far enough to be its own thing. The main idea is that you point it at your models folder and it configures things for you. It reads the GGUF headers, measures how much VRAM you actually have free, and works out context length, GPU offload, CPU/MoE split and KV cache size per model. All of it stays editable per model if you disagree w