相关性不足以为证据:在RAG生成前检测证据缺口
Relevance Is Not Sufficient Evidence: Detecting Evidence Gaps Before Generation in RAG
浏览论文内容
中文总结 AI 辅助
针对RAG中检索证据不足却仍生成答案的问题,提出RINSE方法,在生成前结合三个信号判断证据充分性,在六个数据集上超越现有方法,实现高效本地检测。
中文摘要 AI 辅助
检索增强生成(RAG)将大型语言模型锚定在外部来源上,但检索到的段落往往提及正确的实体,却未提供回答问题所需的事实。即使被指示弃权(不执行),12个生成器仍对40.0%-99.3%的缺乏证据的问题进行了回答。训练生成器弃权(不执行)将决策绑定到模型权重,可能奖励从参数化知识中回忆的答案,并且仍然需要完整的生成器调用。能否在存在任何答案之前,仅从问题和证据判断充分性?我们识别了构建不足证据测试中的陷阱:移除相关证据或将证据与无关问题配对,可能通过词汇重叠或证据位置泄露标签。我们构建了一个配对基准,使用替换、删除和问题交换构造,在控制所选表面特征(如词汇使用)的同时变化答案支持。充分性可以在不生成答案的情况下进行判断,但没有任何单一信号在所有数据集上有效。我们引入了RINSE(相关性不足以为证据),它结合了三个信号:问题的每个部分是否被覆盖,是否有任何段落提供答案,以及一个小型语言模型在共同阅读段落时是否判断它们充分。在六个数据集上,RINSE将充分证据排在不足证据之上,得分为0.837(随机为0.5),超过了10种先前方法中的最佳方法(0.746)和通过API查询的前沿模型(0.784)。其最弱数据集的得分高于任何其他方法的最弱得分(0.684对比0.676)。RINSE在生成前本地运行,在单个GPU上每个问题耗时36.5毫秒。
英文摘要
Retrieval-augmented generation (RAG) grounds large language models in external sources, but retrieved passages often name the right entities without providing the facts needed to answer. Even when instructed to abstain, 12 generators answer 40.0-99.3% of insufficient-evidence questions. Training generators to abstain ties the decision to model weights, may reward answers recalled from parametric knowledge, and still requires a full generator call. Can sufficiency be judged from the question and evidence alone, before any answer exists? We identify pitfalls in constructing insufficient-evidence tests: removing relevant evidence or pairing evidence with unrelated questions can reveal labels through lexical overlap or evidence position. We build a paired benchmark using substitution, deletion, and question-swap constructions that vary answer support while controlling selected surface features, such as word use. Sufficiency can be judged without generating an answer, but no single signal works across all datasets. We introduce RINSE (Relevance Is Not Sufficient Evidence), which combines three signals: whether every part of the question is covered, whether any passage offers an answer, and whether a small language model reading the passages together judges them sufficient. Across six datasets, RINSE ranks sufficient above insufficient evidence with a score of 0.837 (chance 0.5), exceeding the best of 10 prior methods (0.746) and a frontier model queried through an API (0.784). Its weakest dataset scores higher than any other method's weakest (0.684 vs. 0.676). RINSE runs locally before generation, taking 36.5 ms per question on a single GPU.
发表机构
- Northwestern University(西北大学)
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。