arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22228cs.IR

基于提示的弃权(不执行)在误导性语境下失效:对小型冻结RAG模型的对照研究

Prompt-Based Abstention Fails Under Misleading Context: A Controlled Study of Small Frozen RAG Models

Yohanes Andre Setiawan

AI总结:

该研究针对小型冻结RAG模型发现,基于提示的弃权(不执行)无法区分缺失与误导性证据,提出GRAB-RAG基准并测试多种策略,发现其在误导性问题上表现不佳且验证器无法兼顾覆盖率与安全性。

AI中文摘要:

在检索增强生成(RAG)中,缺失证据与误导性证据并非同一问题,但基于提示的弃权(不执行)却将它们同等对待。模型会在缺少上下文时弃权,却不会在上下文具有误导性时弃权。我们推出GRAB-RAG(检索增强生成的分级弃权基准),这是一个配对基准,在《自然问题》(Natural Questions)和《HotpotQA》两个QA基准中,针对四种上下文条件(支持性、降级、缺失和误导性)测试相同问题。在误导性条件下,我们编辑一段黄金段落以支持错误答案,并将其置于其他检索到的段落中。我们在三个小型冻结模型(38亿至80亿参数)上,针对两个QA基准测试了五种弃权策略。模型在证据缺失时会可靠地弃权,但在明确弃权提示下,仍有41.6%的误导性问题被回答,其中63%的回答逐字呼应植入的错误实体。思维链(Chain-of-thought)几乎未提供额外益处。生成器侧冲突检查将该比例降至13.3%,但会丢弃许多正确答案;而NLI验证器恢复了该覆盖率,但当参数记忆与误导性段落对同一错误答案达成一致时,其会失效。基于提示的弃权(不执行)关注上下文是否充足,而非上下文是否正确。两种验证器均无法在不牺牲覆盖率以换取安全性的情况下缩小这一差距。

英文摘要:

Missing and misleading evidence are not the same problem in retrieval-augmented generation (RAG), but prompt-based abstention treats them alike. Models abstain when context is absent, not when it is misleading. We introduce GRAB-RAG (Graded Abstention Benchmark for Retrieval-Augmented Generation), a paired benchmark that tests the same questions across four context conditions (supportive, degraded, missing, and misleading) in Natural Questions and HotpotQA. In the misleading condition, we edit a gold passage to support a wrong answer and place it among other retrieved passages. We test five abstention policies on three small frozen models (3.8B--8B) across two QA benchmarks. Models abstain reliably when evidence is missing, but under explicit abstention prompting still answer 41.6% of misleading questions, with 63% of those answers echoing the planted wrong entity verbatim. Chain-of-thought provides little additional benefit. A generator-side conflict check cuts the rate to 13.3% but discards many correct answers, while an NLI verifier recovers that coverage but fails when parametric memory and the misleading passage agree on the same wrong answer. Prompt-based abstention asks whether context is sufficient, not whether it is correct. Neither verifier closes this gap without trading coverage for safety.

补充信息

↑