arXiv cs.AIOctober 7, 2026
From Probe Scores to Alarm Policies: Operational Validity of Activation Monitors for Language-Model Agents
Excerpt
arXiv:2610.04575v1 Announce Type: cross Abstract: Activation probes can predict safety-relevant properties of language models with high area under the receiver-operating-characteristic curve (AUROC), but deployed agent monitors make thresholded alarm decisions under tight false-alarm budgets. These are different estimands. We introduce an Operational Validity Contract that fixes a monitor's target, observability, identity, timing, intervention unit, comparator, calibration, and cost. We formaliz