arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00585cs.CLcs.IRcs.LG

无充分性的验证:逐块过滤在多跳检索增强生成(RAG)中失效,分解可修复该问题

Verification Without Sufficiency: Per-Chunk Filtering Fails on Multi-Hop RAG, and Decomposition Repairs It

Randhir Kumar

首次发表
浏览论文内容

中文总结 AI 辅助

该研究发现逐块过滤在多跳RAG中失效,经七项控制实验排除其他因素后,提出将验证条件改为分解后的子问题,可显著提升多跳RAG的性能,Qwen2.5-7B分解器已能捕获部分性能提升空间。

中文摘要 AI 辅助

检索增强生成(RAG)的验证通常会对每个检索到的块打分并剔除不合格块。我们证明这对多跳问题无效,并提出可行方案。逐块打分假设单个块是答案的充分前提,但多跳问题的设计逻辑是没有任何单个块能满足这一点,承载答案的段落往往是问题未提及的段落。在HotpotQA、2WikiMultihopQA、MuSiQue数据集上,蕴含评分的AUC分别为0.643、0.523、0.560,而单跳SQuAD数据集上该值为0.951。七项控制变量实验排除了模型容量、前提长度、假设模板、决策阈值、检索器、答案匹配准则、提示的影响。在三个数据集、三种生成器规模、两种提示设置下的端到端实验显示,逐块门控在所有实验组中均显著差于完全不过滤,且其性能损失随生成器能力提升而增大。修复方案是将验证的条件从原始查询改为分解后的子问题。使用MuSiQue的黄金分解时,后续跳数的蕴含评分从随机水平的0.546提升至0.840,配对提升幅度为+0.355,自举区间为[0.331, 0.382]。现成的Qwen2.5-7B分解器在给定问题和顶部检索段落时,达到0.637,捕获了该上限的31%;无检索的分解则达到0.533,低于原始查询对应的水平。迭代检索系统已能生成此类分解,但在验证前会将其丢弃。

英文摘要

Verification for retrieval-augmented generation usually scores each retrieved chunk and drops the ones that fail. We show this cannot work for multi-hop questions, and show what does. Per-chunk scoring assumes one chunk is a sufficient premise for the answer. Multi-hop questions are built so that none is, and the paragraph carrying the answer is the one the question does not name. Entailment scoring reaches 0.643, 0.523 and 0.560 AUC on HotpotQA, 2WikiMultihopQA and MuSiQue, against 0.951 on single-hop SQuAD. Seven controls rule out model capacity, premise length, hypothesis template, decision threshold, retriever, answer-matching criterion and prompt. End to end across three datasets, three generator sizes and two prompts, per-chunk gating is significantly worse than not filtering at all in every cell, and its penalty grows with generator capability. The repair is to condition verification on the decomposed sub-question rather than the original query. Using MuSiQue's gold decomposition, entailment on a later hop rises from 0.546, which is chance, to 0.840, a paired lift of +0.355 with a bootstrap interval of [0.331, 0.382]. An off-the-shelf Qwen2.5-7B decomposer, given the question and the top retrieved paragraph, reaches 0.637 and captures 31% of that ceiling; decomposing without retrieval reaches 0.533, below the original question. Iterative retrieval systems already produce such decompositions and discard them before verifying.

补充信息

↑