← Back to all articles
Reddit r/LocalLLaMASeptember 12, 2026

Agnes-AI/Agnes-3.0-Flash 33B Multimodal, AA score: 36

Excerpt

I find this new model at HF: "Built for demanding work. A 262 144-token context window, adjustable reasoning effort, tool calling, and text, image and video understanding. Architecture Agnes-3.0-Flash is a hybrid-attention decoder: three of every four layers run a gated delta rule (recurrent, with per-layer state independent of sequence length), and the fourth runs standard global attention. Only 18 of the 72 layers therefore hold a KV cache that grows with context." Context length 262 144 token