arXiv cs.AIOctober 7, 2026
EmoRSS: Mitigating Emotion-Induced Over-Refusal in Large Language Models
Excerpt
arXiv:2610.04998v1 Announce Type: new Abstract: Emotional expression can influence the safety decisions of large language models (LLMs), offering a potential avenue for improving safety alignment. Existing studies have mainly focused on how emotional expressions facilitate attacks under harmful requests, while overlooking their effects on benign requests. We find that emotional expression can also systematically increase refusal tendencies on benign requests, leading to unnecessary over-refusal.