arXiv cs.CLOctober 7, 2026
Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge
Excerpt
arXiv:2610.07532v1 Announce Type: cross Abstract: LLMs have advanced rapidly, raising growing concerns about their safety. Recent work has proposed approaches to detect and defend against attacks including defenses at decoding stage that leverage models' hidden states. However, existing decoding-stage defenses suffer from two limitations. First, they introduce a trade-off between safety and over-refusal, where strengthening safety degrades the model's helpfulness on benign queries. Second, many