← Back to all articles
arXiv cs.AIOctober 7, 2026

Benchmarking Jailbreak Guardrails for Embodied Agents

Excerpt

arXiv:2610.06122v1 Announce Type: new Abstract: Embodied agents powered by large language models and vision-language models are increasingly deployed in physical environments, but jailbreak attacks can induce these agents to perform physically harmful actions. A growing number of guardrail methods have been proposed to intercept dangerous behavior before it is executed, yet existing safety benchmarks evaluate the embodied models themselves, leaving it unclear how well these guardrails actually d