← Back to all articles
arXiv cs.AIOctober 7, 2026

The Pushback Paradox: A Two-Probe Diagnostic for Language Model Compliance

Excerpt

arXiv:2610.06673v1 Announce Type: new Abstract: Are language models compliant with user instructions? A model that always complies can be stopped but also exploited, while one that always resists can be neither exploited nor stopped. We contribute an open two-probe benchmark that can place any language model on this spectrum. In the active probe, a user instructs the model to act and accept a lower payoff, which measures exploitability. In the passive probe, the user instructs it to wait and giv