arXiv cs.AIOctober 7, 2026
Target-free Latent Safety Alignment
Excerpt
arXiv:2610.04467v1 Announce Type: cross Abstract: Large language models (LLMs) remain highly vulnerable to jailbreak attacks that induce harmful behaviors and circumvent safety alignment. To defend against such attacks, adversarial training paradigms have been proposed to first simulate failure modes and then train the model to correct them, yielding promising improvements in safety alignment. However, these methods typically construct adversarial samples either by encouraging fixed harmful targ