arXiv cs.AIOctober 7, 2026
CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs
Excerpt
arXiv:2606.11063v2 Announce Type: replace Abstract: AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model. This partially tampers with the untrusted model's trajectory. If the acting model detects such an intervention, it may infer properties of the monitor and adapt to evade the control protocol. We introduce CIAware-Bench, a benchmark for measuring control intervention (CI) awareness across frontier models.