← Back to all articles
Reddit r/LocalLLaMAAugust 27, 2026

Why can't we make MoE routers predict experts needed in the next 5-10 tokens?

Excerpt

Same as title. If we could do that can we potentially do expert caching from ram to vram so it's faster? If not the router itself, can we train a small neural network that predicts the future experts? Sorry if it's a stupid question, I am trying to understand how MoEs work submitted by /u/Hot_Example_4456 [link] [comments]