← Back to all articles
arXiv cs.CLSeptember 21, 2026

CASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation

Excerpt

arXiv:2609.21793v1 Announce Type: cross Abstract: Defenses against jailbreak attacks on Large Language Models (LLMs) operate at different pipeline stages, such as input modification or output guard, but it remains unclear which defenses to deploy at each stage and how to combine them. Prior empirical studies, fragmented by inconsistent attack-success-rate definitions and experimental settings, have evaluated defenses largely in isolation. Here we present the first systematic study, to our knowle