arXiv cs.AIOctober 7, 2026
RAGStress: A controlled benchmark for evaluating retrieval-augmented generation under knowledge-base degradation
Excerpt
arXiv:2610.04691v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is typically evaluated under the implicit assumption that the underlying knowledge base (KB) is clean, leaving the behaviour of RAG systems under realistic KB degradation poorly characterised. We introduce RAGStress, a controlled evaluation benchmark for stress-testing RAG systems under systematic KB corruption. The benchmark pairs four naturalistic corruption types (factual corruption, numeric typo, relevance p