← Back to all articles
arXiv cs.CLSeptember 11, 2026

RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM Safety

Excerpt

arXiv:2609.11758v1 Announce Type: new Abstract: Allowing large language models (LLMs) to retrieve information from a set of trusted documents can increase reliability and reduce hallucination. However, recent work has demonstrated that retrieval-augmented generation (RAG) can have unintended side effects on the overall safety of the generated responses, when prompted for harmful or dangerous content. A clearer understanding of the mechanisms leading to this result is needed, as increasing number