← Back to all articles
arXiv cs.AIOctober 7, 2026

Reactivating Alignment: Defending LLMs from Jailbreaks via Intention-Aware Input-Output Matching

Excerpt

arXiv:2610.04470v1 Announce Type: new Abstract: Large language models (LLMs) remain vulnerable to jailbreak attacks that conceal harmful intent within complex adversarial prompts. Existing defenses primarily rely on input perturbation or harmful-output suppression, but they rarely model where malicious intent resides, resulting in brittle protection and excessive over-refusal. We propose SENTINEL, a plug-and-play, generation-time jailbreak defense that reframes mitigation as an intent extraction