发表机构
Beijing University of Posts and Telecommunications; Humanlaya Data; North China University of Technology(北京邮电大学; 慧览数据; 北方工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出RAG-Stress诊断协议,评估15个系统在三类QA数据集上的表现,发现优先文档的指令误导率更高,且MR随答案位置变化,揭示证据依赖与事实可靠性的差异。
AI 中文摘要
遵循检索到的证据并不能保证事实正确性:误导性证据会诱导模型替换其先前正确给出的答案。标准准确率指标通过将答案替换与已存在的错误相结合,掩盖了这种行为。我们引入RAG-Stress,这是一种受控诊断协议,用于检查检索增强生成中证据依赖的极限。该协议保持问题和参考答案固定,编辑一个断言以支持指定的错误答案,并将两种源优先级策略与证据文本内答案跨度的三个位置交叉。我们在未使用检索时回答正确的问题子集上测量误导率(MR),同时在完整评估集上测量干净准确率。我们评估了15个系统,涵盖API模型、开放模型以及在TriviaQA-RC、HotpotQA和SearchQA上通过强化学习训练的搜索智能体,还额外进行了英语和中文MedQA评估。优先考虑文档的指令始终比允许依赖先验知识的指令产生更高的MR。在三个QA数据集上,跨模型和位置的平均差距为10.9至13.5个百分点。在两种策略下,平均MR遵循“结尾>开头>中间”的顺序,尽管个别模型并不统一遵循此顺序。对500个问题和两个检查点的单独配对审计支持有害覆盖增加,而未建立相应的有益修正的改进。这些发现将证据 adherence 与事实可靠性区分开来,并促使评估检索到的证据是保留、替换还是修正模型的答案。
英文摘要
Following retrieved evidence does not guarantee factual correctness: misleading evidence can induce a model to replace an answer it previously gave correctly. Standard accuracy measures obscure this behavior by combining answer replacement with preexisting errors. We introduce RAG-Stress, a controlled diagnostic protocol for examining the limits of evidence reliance in retrieval-augmented generation. The protocol holds the question and reference answer fixed, edits one assertion to support a designated incorrect answer, and crosses two source priority policies with three positions of the answer span within the evidence text. We measure misleading rate (MR) on each model's subset of questions answered correctly without retrieval, alongside clean accuracy on the full evaluation set. We evaluate fifteen systems spanning API models, open models, and search agents trained with reinforcement learning on TriviaQA-RC, HotpotQA, and SearchQA, with additional English and Chinese MedQA evaluations. Instructions that prioritize documents consistently produce higher MR than those permitting reliance on prior knowledge. Averaged over models and positions, the gap ranges from 10.9 to 13.5 percentage points across the three QA datasets. Mean MR follows End $>$ Beginning $>$ Middle under both policies, although individual models do not uniformly follow this ordering. A separate paired audit of 500 questions and two checkpoints supports increased harmful override without establishing a corresponding improvement in beneficial correction. These findings distinguish evidence adherence from factual reliability and motivate evaluating whether retrieved evidence preserves, replaces, or corrects a model's answers.