← Back to all articles
arXiv cs.LGOctober 2, 2026

Detecting Multi-Agent Collusion Through Multi-Agent Interpretability

Excerpt

arXiv:2604.01151v3 Announce Type: replace-cross Abstract: As LLM agents are increasingly deployed in multi-agent systems, they introduce risks of covert coordination that may evade standard forms of human oversight. While linear probes on model activations have shown promise for detecting deception in single-agent settings, collusion is inherently a multi-agent phenomenon, and the use of internal representations for detecting collusion between agents remains unexplored. We introduce NARCBench, a