arXiv cs.LGOctober 7, 2026
Stepped MoE: Segment-Level Routing with Configurable Inference Complexity
Excerpt
arXiv:2610.07348v1 Announce Type: new Abstract: Training large language models (LLMs) is resource-intensive, and adapting them for diverse deployment scenarios with varying computational constraints remains challenging. While elastic architectures enable flexible model deployment and sparsely activated models allow input-adaptive computation, existing approaches treat these dimensions independently. Moreover, models catered towards on-device edge inference need to conform to the memory and compu