arXiv cs.LGOctober 1, 2026
Advancing Entropy-Level Credit Assignment in RLVR via Proximal Entropy Policy Optimization
Excerpt
arXiv:2609.39402v1 Announce Type: cross Abstract: Value-model-free RLVR methods such as GRPO assign uniform advantages to all tokens in a rollout, ignoring that tokens contribute unequally. Recent methods use token entropy as an importance proxy but compute it globally across the batch, conflating importance with prompt difficulty and positional trends. We argue that importance should instead be measured relative to the local context of each token. We introduce proximal entropy, a local measure