Reddit r/LocalLLaMASeptember 10, 2026
GigaChat-3.5-Reasoning
Excerpt
Hey y'all! We've released a new model in our lineup: GigaChat-3.5 Reasoning. It's a 432B-A28B MoE with Gated DeltaNet for long-context efficiency. We trained domain experts (code, math, general, etc.) with CISPO and then distilled them into a single model via on-policy distillation. In our evals the resulting model lands close to DeepSeek V4 Flash Preview while using 37% fewer tokens in its reasoning traces. Weights are on Hugging Face under MIT: https://huggingface.co/collections/ai-sage/gigach