Reddit r/LocalLLaMASeptember 13, 2026
I built a serverless hosting platform for LoRA adapters with vLLM
Excerpt
It’s always bothered me that after fine-tuning a model for a project, there isn’t a particularly easy way to host it without either running it locally and keeping a GPU on 24/7 or paying for an entire GPU server. There are managed options for LoRA serving on top of vLLM (AWS), but you generally still end up paying for an entire instance. I started wondering: if 99%+ of the model weights are identical between the base model and something like a rank 8–32 LoRA/QLoRA adapter, why does each adapter