arXiv cs.AIOctober 7, 2026
AgentPrivArena: Evaluating and Auditing Real-world AI Agent Privacy
Excerpt
arXiv:2610.06454v1 Announce Type: new Abstract: The rapid advancement of LLM agents has enabled systems to autonomously perform complex tasks through external tools, but their growing access to personal data introduces significant privacy risks. Existing benchmarks primarily evaluate LLM agent privacy through simulated trajectories and outcome-based metrics, limiting their ability to capture privacy risks arising during multi-step agent execution. In this work, we introduce AgentPrivArena, a fra