arXiv cs.LGOctober 7, 2026
Principled Under Pressure: Post-Training Decides Whether LLMs Act on Their Own Moral Judgment
Excerpt
arXiv:2610.08670v1 Announce Type: new Abstract: Language models increasingly act as agents. An agent that says an action is wrong and then takes it anyway is a different failure from one that does not know better, and evaluations of stated values cannot see it. We build a pre-registered panel of 248 scenarios across five kinds of pressure. Each scenario is posed twice to the same model, once as the agent choosing what to do and once in the third person asking which option is right, so the model'