← Back to all articles
arXiv cs.AIOctober 7, 2026

Evaluating System One Models for Agent Security Decisions: Reliability, Calibration, and Selective Automation

Excerpt

arXiv:2609.33401v3 Announce Type: replace-cross Abstract: Agentic software connects language models to tools that modify files and interact with external services. Developers use model-based security checks to screen external content and user requests before agents act. System One models answer developer-defined questions with probabilities over predefined answers such as safe or unsafe, but classification accuracy alone does not establish whether these probabilities support reliable automation.