← Back to all articles
arXiv cs.LGOctober 1, 2026

Training LLM Judges from Language Feedback via Position-Selective Self-Distillation

Excerpt

arXiv:2609.38792v1 Announce Type: cross Abstract: We study training LLM judges from natural language feedback, especially for subjective tasks where the verdict depends strongly on which evaluation criteria the judge invokes and how it weighs them. The dominant approach, outcome-supervised RL (e.g., GRPO), credits every token in the rollout with a single scalar determined only by the accuracy of the final verdict, providing no separate credit at the criterion-choice tokens and ignoring the rich