← Back to all articles
arXiv cs.LGOctober 2, 2026

Why GRPO Needs Normalization: A Local-Curvature Perspective on Adaptive Gradients

Excerpt

arXiv:2601.23135v2 Announce Type: replace Abstract: Reinforcement learning (RL) has become a key driver of language model reasoning. Among RL algorithms, Group Relative Policy Optimization (GRPO) is the de facto standard, avoiding the need for a critic by using per-prompt baselines and variance normalization. Yet why and when this normalization helps remains unclear. In this work, we answer both questions through the lens of local curvature of the sequence-level policy gradient: standard deviati