Reddit r/LocalLLaMASeptember 7, 2026
Qwen3.5 0.8B on CPU
Excerpt
Since the Qwen3.5 0.8B model is an interesting one for small specialized fine tunes, I was curious how fast it can run on CPUs. Why CPUs? Mainly because I want to use it as a local dictation cleanup model when I'm using the GPU for something else. Over the weekend, I let Codex build a small C++ engine and a custom 4-bit format, H128/Q4-G32-DOT4 , with activation-based calibration and blockwise error compensation. The resulting model has a 425 MB weight payload , roughly 71 MB smaller than Unslot