arXiv cs.LGOctober 2, 2026
Why GRPO Needs Normalization: A Local-Curvature Perspective on Adaptive Gradients
Excerpt
arXiv:2601.23135v2 Announce Type: replace Abstract: Reinforcement learning (RL) has become a key driver of language model reasoning. Among RL algorithms, Group Relative Policy Optimization (GRPO) is the de facto standard, avoiding the need for a critic by using per-prompt baselines and variance normalization. Yet why and when this normalization helps remains unclear. In this work, we answer both questions through the lens of local curvature of the sequence-level policy gradient: standard deviati