arXiv cs.AIOctober 7, 2026
BRANCH-MoE: Balance-Aware Tree Routing for Large Embedding Models
Excerpt
arXiv:2610.06725v1 Announce Type: cross Abstract: Mixture-of-experts (MoE) layers increase model capacity without a proportional increase in per-example computation. However, conventional flat routers can yield imbalanced expert utilization and treat experts as an unstructured collection, whose indices carry no topological meaning. We introduce {\bf BRANCH-MoE}, a routing architecture that places \(E\) experts at the leaves of a binary decision tree of depth \(\log_2 E\). At each internal node t