arXiv cs.AIOctober 7, 2026
Evaluating System One Models for Agent Security Decisions: Reliability, Calibration, and Selective Automation
Excerpt
arXiv:2609.33401v3 Announce Type: replace-cross Abstract: Agentic software connects language models to tools that modify files and interact with external services. Developers use model-based security checks to screen external content and user requests before agents act. System One models answer developer-defined questions with probabilities over predefined answers such as safe or unsafe, but classification accuracy alone does not establish whether these probabilities support reliable automation.