arXiv cs.LGOctober 2, 2026
Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
Excerpt
arXiv:2610.00574v1 Announce Type: new Abstract: Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout groups, but different objectives can still exhibit uneven learning progress. We study this behavior through advantage energy, the sum of a reward's squared advantages over a batch. Under idealized GDPO normalization, we sho