← Back to all articles
Reddit r/LocalLLaMASeptember 18, 2026

RTX 5090 Bonsai 2 27B vs Gemma 4 12B vs Qwen 3.5 9B Japanese voxel pagoda

Excerpt

Yesterday I saw that prismml just dropped the new bonsai which is a heavily-quantized version of qwen 3.8 27b claiming over 98% top-1% comparing to fp16, while weighing from ~6gb at Q1 to ~8gb at Q2 . I wanted to see how good this model is compared to other models in this memory range and my choice fell on gemma 4 12B and qwen 3.5 9B I ran all the tests on an rtx 5090 and gave the models an identical pagoda prompt(the results are one shot, and yes, bonsai used a lot more tokens, but 90% of the b