arXiv cs.AIOctober 7, 2026
Boosting Transferable Adversarial Attacks against Deep Reinforcement Learning
Excerpt
arXiv:2610.06083v1 Announce Type: cross Abstract: Most adversarial attacks on deep reinforcement learning (DRL) assume white-box access to the victim policy, which rarely holds in practice. This paper studies transfer-based black-box attacks on DRL: the attacker crafts observation perturbations on a white-box surrogate agent and feeds them to an unknown victim. We formulate the attack as return minimization under a per-step perturbation budget. We first show that transplanting transferable image