arXiv cs.LGAugust 17, 2026
Localization then Neutralization: Gradient-guided Token Suppression against Visual Prompt Injection Attack
Excerpt
arXiv:2605.25194v2 Announce Type: replace Abstract: Adversarial images pose a severe security threat to multimodal large language models through prompt injection. Existing defenses largely lack a principled understanding of the underlying mechanisms and struggle to balance efficiency and defense utility. In this work, we show that successful adversarial attacks do not rely on the entire image uniformly but instead depend on a small subset of critical image tokens. Based on this insight, we propo