arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RAGStress:一个用于评估知识库退化下检索增强生成的受控基准

RAGStress: A controlled benchmark for evaluating retrieval-augmented generation under knowledge-base degradation

Shiqi Yang, Jiekai Ma, Gaoyuan Du

arXiv 2610.04691首次发表:更新:

发表机构

Amazon(亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RAGStress是一个受控基准,通过四种损坏类型和三种严重程度系统性破坏知识库,评估RAG在退化下的鲁棒性,发现清洁检索掩盖差异且语义损坏危害更大。

AI 中文摘要

检索增强生成(RAG)通常在知识库(KB)清洁的隐含假设下进行评估,导致RAG系统在现实知识库退化情况下的行为特征化不足。我们引入RAGStress,一个用于在系统性知识库损坏下对RAG系统进行压力测试的受控评估基准。该基准将四种自然主义损坏类型(事实损坏、数字笔误、相关性污染和矛盾注入)与三种严重程度(细微、中等和明显)配对,采用基于57个MMLU主题和182,546个文档的单知识库、元数据过滤实验设计。在52,500次模型-问题-条件评估中,RAGStress揭示:清洁检索可能掩盖鲁棒性差异;语义保真度损坏比信号效用扰动危害大得多;无检索准确率不能预测损坏检索的鲁棒性;混合知识库准确率不应被视为最坏情况鲁棒性。我们记录了基准的预期用途、支持的声明和局限性,并提供包含生成脚本、损坏提示、元数据模式和评估代码的工件包。RAGStress旨在作为知识库损坏下RAG鲁棒性的受控压力测试,而非通用模型排行榜。

英文摘要

Retrieval-Augmented Generation (RAG) is typically evaluated under the implicit assumption that the underlying knowledge base (KB) is clean, leaving the behaviour of RAG systems under realistic KB degradation poorly characterised. We introduce RAGStress, a controlled evaluation benchmark for stress-testing RAG systems under systematic KB corruption. The benchmark pairs four naturalistic corruption types (factual corruption, numeric typo, relevance poisoning, and contradiction injection) with three severity levels (subtle, moderate, and obvious) over a single-KB, metadata-filtered experimental design built from 57 MMLU subjects and 182,546 documents. Across 52,500 model-question-condition evaluations, RAGStress reveals that clean retrieval can mask robustness differences, semantic-fidelity corruptions are substantially more harmful than signal-utility perturbations, no-retrieval accuracy does not predict corrupted-retrieval robustness, and mixed-KB accuracy should not be treated as worst-case robustness. We document the benchmark's intended use, supported claims, and limitations, and provide an artifact bundle including generation scripts, corruption prompts, metadata schema, and evaluation code. RAGStress is intended as a controlled stress test for RAG robustness under KB corruption, not as a general model leaderboard.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑