arXiv cs.LGOctober 1, 2026
Training LLM Judges from Language Feedback via Position-Selective Self-Distillation
Excerpt
arXiv:2609.38792v1 Announce Type: cross Abstract: We study training LLM judges from natural language feedback, especially for subjective tasks where the verdict depends strongly on which evaluation criteria the judge invokes and how it weighs them. The dominant approach, outcome-supervised RL (e.g., GRPO), credits every token in the rollout with a single scalar determined only by the accuracy of the final verdict, providing no separate credit at the criterion-choice tokens and ignoring the rich