← Back to all articles
Reddit r/LocalLLaMAAugust 22, 2026

Qwen 3.5 4B IQ2_XS: +16.67% Reasoning Performance From Tensor-Level Allocation

Excerpt

I was finally able to replicate tensor level allocation outside the Gemma family. https://huggingface.co/ByteOtter/Qwen3.5-4B-CADA-IQ2_XS After the Gemma 4 12b, e4b and gemma 3 4b results, I attempted to expand into qwen and ran into a few walls. After 2 version updates and a slightly different approach, I was able to replicate the effect on Qwen. The result: BF16 reasoning: 78.125 Stock IQ2_XS + imatrix: 46.875 QLAB allocation + same imatrix: 54.688 That's +7.812 percentage points, or a +16.67%