← Back to all articles
arXiv cs.CLSeptember 9, 2026

Membrane: A Self-Evolving Contrastive Safety Memory for LLM Agent Defense

Excerpt

arXiv:2606.05743v2 Announce Type: replace-cross Abstract: Despite advances in safety alignment, large language models remain vulnerable to continuously evolving jailbreaks. Existing fine-tuned safety classifiers cannot adapt to these evolving attacks, while adaptive memory-based guardrails tend to over-refuse benign queries that resemble stored attacks. We propose Membrane, a self-evolving guardrail built on Contrastive Safety Memory (CSM): each cell pairs the conditions for blocking a harmful q