← Back to all articles
arXiv cs.LGOctober 2, 2026

Range-GRPO: Policy Optimization via Pairwise Relations among Reward Intervals

Excerpt

arXiv:2610.01548v1 Announce Type: new Abstract: As the use of large language models (LLMs) expands, post-training has become increasingly important for adapting them to downstream tasks. However, obtaining reliable supervision remains costly, especially in domains without reference answers or executable verifiers. LLM-as-a-Judge provides scalable pseudo-rewards for unlabeled responses, but a single point score does not explicitly represent reward uncertainty. This motivates representing pseudo-r