Reddit r/LocalLLaMAAugust 27, 2026
pareto frontiers
Excerpt
Qwen models are both pareto frontiers in total size AND active parameters size of all open weights models so far. If this trend continues, we might see sparser and more capable models really soon, given that this is a preview of Qwen4 and is probably undertrained. What do you people think? Will the trend continue? Will Qwen stay in the lead? And more importantly, does it scale up? (e.g. would Qwen4 architecture at larger scales be even better? or diminishing returns?) I personally like the path