← Back to all articles
arXiv cs.AIOctober 7, 2026

RAGStress: A controlled benchmark for evaluating retrieval-augmented generation under knowledge-base degradation

Excerpt

arXiv:2610.04691v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is typically evaluated under the implicit assumption that the underlying knowledge base (KB) is clean, leaving the behaviour of RAG systems under realistic KB degradation poorly characterised. We introduce RAGStress, a controlled evaluation benchmark for stress-testing RAG systems under systematic KB corruption. The benchmark pairs four naturalistic corruption types (factual corruption, numeric typo, relevance p