← Back to all articles
arXiv cs.LGOctober 1, 2026

Advancing Entropy-Level Credit Assignment in RLVR via Proximal Entropy Policy Optimization

Excerpt

arXiv:2609.39402v1 Announce Type: cross Abstract: Value-model-free RLVR methods such as GRPO assign uniform advantages to all tokens in a rollout, ignoring that tokens contribute unequally. Recent methods use token entropy as an importance proxy but compute it globally across the batch, conflating importance with prompt difficulty and positional trends. We argue that importance should instead be measured relative to the local context of each token. We introduce proximal entropy, a local measure