← Back to all articles
arXiv cs.AIOctober 7, 2026

Agent MechSuits: Mechanistic Subspace Safety Steering for Multi-Turn CLI Agents

Excerpt

arXiv:2606.22673v2 Announce Type: replace Abstract: Command-Line Interface (CLI) agents based on large language models (LLMs) demonstrate remarkable autonomous capabilities, but they also introduce significant safety and misuse risks during multi-turn interactions with external environments. Existing safety mechanisms mainly rely on external guardrails, which have a limited ability to perform fine-grained behavioral control during execution. Meanwhile, recent mechanistic interpretability methods