arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估RAG中证据感知检索的下游效用

Assessing the Downstream Utility of Evidence-Aware Retrieval in RAG

Utshab Kumar Ghosh, Debayan Mukhopadhyay, Shubham Chatterjee

arXiv 2608.26379首次发表:更新:

发表机构

Missouri University of Science and Technology; University of Calcutta(密苏里科技大学; 加尔各答大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究在多基准及TREC RAG 2025设置中,考察证据感知检索的下游效用,发现其价值不统一,需针对特定用途评估RAG评估方法。

AI 中文摘要

检索增强生成(Retrieval-Augmented Generation, RAG)的检索评估越来越围绕检索到的段落是否包含可支持生成的证据,而非仅主题相关性展开。本研究探讨这种与下游证据需求更紧密的对齐是否也会使检索评估对基于其的决策更有用。在五个检索基准和端到端TREC RAG 2025设置中,我们从四个角色考察答案支持信号:比较检索器、指导检索训练和系统选择、预测下游答案质量、过滤提供给生成器的证据。该信号改变了检索排名,但其下游价值并不统一:它无法可靠改进检索器训练;将其用于系统选择的益处取决于生成器被指示如何使用检索到的证据;基于它的检索分数无法稳健预测未见过主题的答案质量。在直接证据干预中,人工标注者确认过滤优先保留包含有用答案证据的段落,但不同答案评估者对生成的答案是否改善得出不同结论。这些结果表明,使检索评估更紧密反映生成所需的证据本身,并不一定使该评估的每项下游使用更可靠。因此,应针对RAG评估方法旨在支持的特定比较、决策和结论对其进行评估。

英文摘要

Retrieval evaluation for retrieval-augmented generation (RAG) is increasingly designed around whether retrieved passages contain evidence that can support generation, rather than topical relevance alone. We study whether this closer alignment with downstream evidence needs also makes retrieval evaluation more useful for the decisions built from it. Across five retrieval benchmarks and an end-to-end TREC RAG 2025 setting, we examine an answer-support signal in four roles: comparing retrievers, guiding retrieval training and system selection, predicting downstream answer quality, and filtering the evidence supplied to a generator. The signal changes retrieval rankings, but its downstream value is not uniform. It does not reliably improve retriever training; the benefit of using it for system selection depends on how the generator is instructed to use the retrieved evidence; and retrieval scores based on it do not robustly predict answer quality on unseen topics. In a direct evidence intervention, human annotators confirm that filtering preferentially preserves passages containing useful answer evidence, yet different answer evaluators reach different conclusions about whether the resulting answers improve. These results show that making retrieval evaluation more closely reflect the evidence needed for generation does not by itself make every downstream use of that evaluation more reliable. RAG evaluation methods should therefore be assessed with respect to the particular comparisons, decisions, and conclusions they are intended to support.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑