← Back to all articles
arXiv cs.CLSeptember 22, 2026

Measuring the Assistant's Harmlessness Preferences on the User Turn

Excerpt

arXiv:2609.23935v1 Announce Type: new Abstract: Post-training turns a general next-token predictor into a chat model with a persistent assistant persona. If that persona is a character the model plays only on its own turns, its preferences should govern what the assistant says, not what the model predicts other speakers will say. We test this boundary and find that it does not hold: a safety-relevant preference of the assistant---for harmless over harmful tasks---shapes the model's predictions e