← Back to all articles
arXiv cs.AIOctober 7, 2026

EmoRSS: Mitigating Emotion-Induced Over-Refusal in Large Language Models

Excerpt

arXiv:2610.04998v1 Announce Type: new Abstract: Emotional expression can influence the safety decisions of large language models (LLMs), offering a potential avenue for improving safety alignment. Existing studies have mainly focused on how emotional expressions facilitate attacks under harmful requests, while overlooking their effects on benign requests. We find that emotional expression can also systematically increase refusal tendencies on benign requests, leading to unnecessary over-refusal.