← Back to all articles
Reddit r/LocalLLaMASeptember 19, 2026

Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro

Excerpt

Meet Inco Splash, open-source inference engine, built around the model and around Apple silicon. Up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents. Requirements: M3 or newer, macOS 26.4+, 36 GB Get started with a single command: brew install incoai/tap/splash splash serve --model incoai/Qwen3.8-27B-Splash That is the whole setup. Point your agent at it, works with Claude Code, OpenCode, Codex, or Hermes Prefer an app? Also, available in LM Studio