arXiv cs.CLOctober 7, 2026
Beyond Refusal Patterns: Safe-Role Internalization for Robust and Generalizable LLM Safety Alignment
Excerpt
arXiv:2610.07023v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved remarkable capabilities but remain vulnerable to jailbreak attacks that elicit harmful or unsafe outputs. Existing safety alignment approaches, including Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), often require substantial attack-specific supervision and computational resources, while remaining susceptible to shallow safety alignment and over-refusal. To address t