arXiv cs.AIOctober 7, 2026
A Bird's-Eye View of Iterative Reward Design
Excerpt
arXiv:2610.04364v1 Announce Type: cross Abstract: Designing effective reward functions in RL typically requires substantial expertise and trial and error. Recent work automates this process with LLM-based systems that generate and iteratively improve reward code using policy feedback. However, these methods are often hard to compare because they differ in implementation details, feedback assumptions, and evaluation environments. To address this, we introduce a Benchmark for Iterative Reward Desi