← Back to all articles
Reddit r/LocalLLaMAAugust 20, 2026

Ornith-1.5-35B-A3B Q4 running 60tk/s on 4070Ti.

Excerpt

Not everyone has the disposable income to build a small data center, so making this post for the underdogs as I was very surprised by the performance/results of this 35B MOE model. Context is admittedly tight and will get laughed at by the big boys. I tried to think of something inspirational to say here but failed, so you just get laughed at. Sorry. The thought here is to push as many active experts into Vram and offload the rest into system ram. At 27 it leaves about 1gig of overhead for KV ca