arXiv cs.AIOctober 7, 2026
Is Escalation Worth It? On the Depth of LLM Cascades
Excerpt
arXiv:2605.06350v2 Announce Type: replace-cross Abstract: LLM cascades, in which a cheap model defers to an expensive one on low-confidence queries, are widely used to reduce inference cost. Given a pool of models, a practitioner must decide how many models to include and where to set each deferral threshold. We derive first-order optimality conditions showing that, at an optimum, the ratio of expected accuracy gain to expected downstream cost is equal across deferral boundaries. A local search