← Back to all articles
arXiv cs.AIOctober 7, 2026

CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs

Excerpt

arXiv:2606.11063v2 Announce Type: replace Abstract: AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model. This partially tampers with the untrusted model's trajectory. If the acting model detects such an intervention, it may infer properties of the monitor and adapt to evade the control protocol. We introduce CIAware-Bench, a benchmark for measuring control intervention (CI) awareness across frontier models.