← Back to all articles
Reddit r/LocalLLaMASeptember 21, 2026

[Splash Engine] Qwen3.8-27B in native 8-bit at 37–55 tok/s on Apple Silicon: Extending Splash to Q8, 256k context scaling, and the "Reasoning Cliff"

Excerpt

https://preview.redd.it/nulsv53o8vqh1.png?width=4500&format=png&auto=webp&s=74765dbd409f4c221640f9f6000a685f6fdbb242 Spent weekend benchmarking the Splash engine (by Incoai) and extending its architecture to native 8-bit on Apple Silicon (M5 Pro, 64 GB unified memory). Splash is a compiled C++ and Metal speculative decoding engine designed specifically for Apple Silicon. Upstream Splash pioneered a blisteringly fast speculative decoding pipeline for 4-bit models (~60 tok/s). However, aggressive