Reddit r/LocalLLaMASeptember 19, 2026
M1 Max 32GB, trying to run Qwen 3.8 27B at decent speeds and context
Excerpt
I would like to stay at q4, and serve the model to various coding harnesses. None of the models or tools I have tried will run without running out of memory. I am hoping for some advice. oMLX? MTPLX? llama.cpp? Has anyone gotten a good Qwen 3.8 27B running ok on an M1 Max 32GB MacBook Pro? Edit: Decent speeds means it's not all the way down at 1 t/s. And if anyone has other suggestions to get close to the quality but with another model, I'd love suggestions. submitted by /u/greggh [link] [commen