arXiv cs.LGOctober 1, 2026
CATCH: A Controllable Analysis Testbed for Reward Hacking in Coding RL
Excerpt
arXiv:2609.39533v1 Announce Type: cross Abstract: During reinforcement learning with verifiable rewards (RLVR), large language models (LLMs) can exploit loopholes in their environments to obtain high rewards without improving the intended capabilities, i.e., reward hacking. Despite its risks to training efficiency and safety, monitoring and mitigating reward hacking during training remain challenging, which is limited by a lack of testbeds that reproduce hacking and reliably identify it. We intr