← Back to all articles
Reddit r/LocalLLaMAAugust 27, 2026

llama : add --n-cpu-ffn option by John-194 · Pull Request #26622 · ggml-org/llama.cpp

Excerpt

tl;dr faster dense models for low VRAM people option similar to the existing --n-cpu-moe It puts user specified amount of FFN sublayers for dense models. PR by u/Stainless-Bacon submitted by /u/jacek2023 [link] [comments]