Reddit r/LocalLLaMASeptember 23, 2026
M5U base 96GB inference numbers for Q3.8FN after 112M tokens
Excerpt
TLDR; Base M5 Ultra 96 GB ran Q3.8 FN aggregate 3.2k PP and ~170 TG in 4 concurrency Alert: Numbers and custom server details at end are AI assisted So the good news is that I got the base model on launch day with only 64 core GPU. All benchmarks are for current maxed out model, which looks very sweet but it is 2x the price. Plus, the estimated delivery date is in 2027! So I thought about maximizing current device on hand and share some of the ideas that worked for me. I did see some odd sub age