Reddit r/LocalLLaMASeptember 12, 2026
Nex-N2.5-mini-MLX-4bit on Apple M5 Max — 133.6 tok/s — llm-bench.io
Excerpt
Another new model dropped in the course of this week that is well deployable on consumer hardware: Nex N2.5 Mini I went with the recommended settings for the best generation quality and ran a few benchmarks: temperature: 0.7 top_p: 0.95 top_k: 40 reasoning_effort: high I must say, the outcome is not bad at all - really good generation speed and prompt processing, okay memory footprint and good quality across the board. Will for sure give it a try to fuel my agents and might also try to do some c