arXiv cs.AIOctober 7, 2026
AI Safety via Debate is Compromised by Cognitive Biases
Excerpt
arXiv:2610.05461v1 Announce Type: new Abstract: Reinforcement learning from human feedback (RLHF) has played a central role in making large language models responsive to human instructions. However, human evaluators often favor flattering or persuasive responses over truthful ones, creating incentives for models to appeal to evaluators at the expense of accuracy. AI safety via debate has been proposed as a way to improve the supervision of language models: in this paradigm, two agents argue oppo