arXiv cs.LGOctober 2, 2026
Detecting Multi-Agent Collusion Through Multi-Agent Interpretability
Excerpt
arXiv:2604.01151v3 Announce Type: replace-cross Abstract: As LLM agents are increasingly deployed in multi-agent systems, they introduce risks of covert coordination that may evade standard forms of human oversight. While linear probes on model activations have shown promise for detecting deception in single-agent settings, collusion is inherently a multi-agent phenomenon, and the use of internal representations for detecting collusion between agents remains unexplored. We introduce NARCBench, a