← Back to all articles
arXiv cs.CLOctober 7, 2026

Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage?

Excerpt

arXiv:2606.05647v2 Announce Type: replace-cross Abstract: AI coding agents are increasingly embedded in real-world software development, collaborating with human developers while gaining broader access to codebases and tools. This creates a new attack surface: an agent can exploit human trust to sabotage development, for instance by inserting malicious code to accomplish a hidden side task. Most prior work studies AI sabotage in AI-only settings, paying limited attention to the role of human ove