arXiv cs.LGOctober 1, 2026
Semifactual Credit-Augmented Policy Optimization
Excerpt
arXiv:2609.40360v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions remain sensitive to task-irrelevant prompt features. We investigate this sensitivity through semifactual prompt interventions that preserve the underlying problem and its answer. Our analysis reveals substantial variation in token-level sensitivity and shows that suppressing high-drift token candidates