← Back to all articles
Reddit r/LocalLLaMAAugust 20, 2026

AirLLM - Recent Updates - with Qwen3.8-27B, Kimi-K3 too

Excerpt

AirLLM dramatically reduces inference memory usage, letting 70B large language models run on a single 4GB GPU card — without quantization, distillation, or pruning. You can even run 405B Llama 3.1 on 8GB , DeepSeek-V3 (671B) on ~12GB , and Kimi K3 (2.8T) — the largest open-source model released to date — on under 4GB , because sparse MoE models stream one expert at a time rather than a whole layer. Updates [ 2026/08 ] Qwen3.8-27B support: Qwen's new dense VL (Gated DeltaNet + Gated Attention, na