← Back to all articles
arXiv cs.AIOctober 7, 2026

Boosting Transferable Adversarial Attacks against Deep Reinforcement Learning

Excerpt

arXiv:2610.06083v1 Announce Type: cross Abstract: Most adversarial attacks on deep reinforcement learning (DRL) assume white-box access to the victim policy, which rarely holds in practice. This paper studies transfer-based black-box attacks on DRL: the attacker crafts observation perturbations on a white-box surrogate agent and feeds them to an unknown victim. We formulate the attack as return minimization under a per-step perturbation budget. We first show that transplanting transferable image