Reddit r/LocalLLaMASeptember 1, 2026
New Model: Spark-X2.5-4B, Spark-X2.5-1.7B
Excerpt
I was browsing HF for small LLMs and run into this model. It does not seem to be a fine tune - the model has its own architecture. https://huggingface.co/XHToken/Spark-X2.5-1.7B https://huggingface.co/XHToken/Spark-X2.5-4B There are 4B/1.7B versions - the benchmark is quite interesting (4B is neck and neck with Qwen 3.5 9B). The HF page claims both models support native 1M context size . Currently does not run out of the box on llama.cpp - pending this PR: https://github.com/ggml-org/llama.cpp/p