arXiv cs.AIOctober 2, 2026
False Floors: LLM Safety Routing Evaluations Break Under Distribution Shift
Excerpt
arXiv:2610.01535v1 Announce Type: cross Abstract: Safety routers send each request to one of several models and are judged against the best single model. A major routing benchmark picks that comparator on the evaluation data. In the benchmark's own setting this is harmless, but under distribution shift it is not. On HELM Safety the selection cost is 0.003-0.030 of harm under random splits and 0.045-0.113 under held-out categories, comparable to the whole deficit attributed to routing, with its d