arXiv cs.AIOctober 7, 2026
Low-Resource Safety Failures Are Action Failures, Not Representation Failures
Excerpt
arXiv:2606.01196v2 Announce Type: replace-cross Abstract: Language models often answer harmful requests in low-resource languages (LRLs) that they refuse in high-resource languages (HRLs). Across three instruction-tuned models and 23 languages, harmful refusal falls from 87.9% in HRLs to 43.9% in LRLs, while harmless refusal remains low. A common explanation is that models represent harmfulness weakly in LRLs. We test whether harmfulness is instead represented but does not reliably produce refus