Reddit r/LocalLLaMAAugust 27, 2026
N-gram vs Experts explained
Excerpt
Since Qwen's dropped the Qwen4Exp architecture bomb that focus on offloading parameters to n-gram instead of pure mixture of experts, I dug into this and learned quite a lot. Here's the summary. Expect mistakes from human's writing lol. TLDR: MoEs do reasoning, N-grams do recalling. At the current tech frontier, N-gram can offload upto ~25% weight before losing advantage and we get most benefit from using SSD to store those instead of RAM. So 176B models became 125 (RAM) +51B (SSD) instead. Long