← Back to all articles
Reddit r/MachineLearningSeptember 27, 2026

A learned LLM router scored 0.84 AUC. Shuffling the labels within each task still scored 0.838 [R]

Excerpt

I trained a router to choose between a cheap and an expensive model. It scored 0.84 AUC held out. Then I shuffled the labels within each task — preserving each task's escalation rate, destroying all per-item signal — retrained, and got 0.838. It had learned to recognize the task, not the difficulty. Eliminated in turn: too few labels (110k from RouteLLM's released set) the architecture (linear probe, similarity-weighted ranking, fine-tuned encoder) the representation (a probe recovers human diff