← Back to all articles
arXiv cs.LGOctober 2, 2026

Constitutional Value Potentials: reading and steering internal priority margins in language models

Excerpt

arXiv:2606.15420v2 Announce Type: replace Abstract: Safety evaluations test a policy on prompts that omit the incentive information deployment supplies: a commission, a performance score, a dashboard naming which action pays best. We measure what that omission hides. In MoneyWorld, a synthetic workplace environment, we train five instruction-tuned models from three families with RL on non-safety tasks in which a visible payoff signal identifies a rewarded shortcut that sacrifices task quality. W