← Back to all articles
arXiv cs.LGAugust 17, 2026

Localization then Neutralization: Gradient-guided Token Suppression against Visual Prompt Injection Attack

Excerpt

arXiv:2605.25194v2 Announce Type: replace Abstract: Adversarial images pose a severe security threat to multimodal large language models through prompt injection. Existing defenses largely lack a principled understanding of the underlying mechanisms and struggle to balance efficiency and defense utility. In this work, we show that successful adversarial attacks do not rely on the entire image uniformly but instead depend on a small subset of critical image tokens. Based on this insight, we propo