← Back to all articles
Reddit r/LocalLLaMASeptember 17, 2026

Ternary Bonsai 2 (27B) just released on Hugging Face. At <6GB in size, it can even run locally in-browser on WebGPU.

Excerpt

The model is derived from Qwen3.8-27B, a 27B hybrid-attention causal language model (architecture unchanged), but uses ternary weights to shrink model size down to - Collection: https://huggingface.co/collections/prism-ml/bonsai-2 - Demo: https://huggingface.co/spaces/webml-community/ternary-bonsai-2-webgpu-kernels submitted by /u/xenovatech [link] [comments]