← Back to all articles
arXiv cs.LGOctober 1, 2026

Semifactual Credit-Augmented Policy Optimization

Excerpt

arXiv:2609.40360v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions remain sensitive to task-irrelevant prompt features. We investigate this sensitivity through semifactual prompt interventions that preserve the underlying problem and its answer. Our analysis reveals substantial variation in token-level sensitivity and shows that suppressing high-drift token candidates