← Back to all articles
arXiv cs.CLAugust 17, 2026

Understanding and Mitigating Over-refusal for Large Language Models via Representation Intervention

Excerpt

arXiv:2511.19009v2 Announce Type: replace-cross Abstract: Large language models (LLMs) demonstrate powerful capabilities across various natural language processing tasks,yet their inherent safety vulnerabilities undermine the reliable application of LLMs in real-world scenarios. To enhance LLM safety, various jailbreak defense methods have been proposed to guard against harmful outputs. However, improvements in model safety often come at the cost of severe over-refusal, failing to strike a good