← Back to all articles
Reddit r/LocalLLaMASeptember 24, 2026

model : add Ling 3.0 VL support by aetherbird · Pull Request #29151 · ggml-org/llama.cpp

Excerpt

Model Overview Ling-3.0-flash-VL inherits the language, reasoning, and long-context capabilities of Ling-3.0-flash, while extending them with native image and video understanding. The model has 124B total parameters, with only 5.5B parameters activated per token, and supports a context window of up to 256K tokens. The architecture of Ling-3.0-flash-VL is designed to integrate visual information into real-world reasoning and agentic workflows. A ViT visual encoder extracts features from images an