arXiv cs.AIAugust 18, 2026
RLCascadeRouter: Quality-Estimator-Free Cascade Routing via Reinforcement Learning
Excerpt
arXiv:2608.15817v1 Announce Type: new Abstract: The growing ecosystem of large language models (LLMs) offers huge potential to optimize performance-cost trade-offs. However, their heterogeneous capabilities and inference costs make efficiently routing queries a significant challenge. Existing paradigms are inflexible: one-shot routers commit before observing responses, whereas conventional cascades stop adaptively but follow a fixed model order. Cascade routing removes both restrictions by recon