arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13956cs.IR

检索器冗余性与多样性如何影响检索增强生成(RAG)的有效性

How retriever redundancy and diversity impact RAG effectiveness

Jonathan J Ross, Bevan Koopman, Anton van der Vegt, Guido Zuccon

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探究RAG中检索文档的冗余性与多样性对答案正确性的影响,发现重复冗余与LLM转述无显著增益,而多样文档可将正确性提升17%-47%,并揭示其提升源于体裁多样性,为RAG检索方法优化提供方向。

中文摘要 AI 辅助

在检索增强生成(RAG)中,检索器通常会根据文档与查询的个体相关性对文档进行排序,而生成器则会基于检索到的文档整体生成答案。本文研究了检索到的文档集的冗余性与多样性如何从答案正确性的角度影响生成器。以往研究的结果混杂:部分研究表明冗余性通过强化相关信息提升生成效果,另一些研究则表明对同一内容的大型语言模型(LLM)转述可能有益。许多此类研究未控制混淆因素,例如文档是否包含确切答案、参数知识是否发挥作用。我们开展了控制严格的实验,研究检索到的文档集的三种关键场景:1)重复场景(同一文档的精确副本);2)转述场景(同一文档的LLM转述版本);3)多样场景(来自不同体裁、各以不同形式包含相关信息的文档)。我们控制文档是否包含确切匹配或转述形式的答案,使用FictionalQA进行评估——这是一个合成的虚构问答数据集,确保LLM生成器的先验知识无法回答该问题,答案必须来自检索到的文档。结果显示,重复冗余与LLM转述均未显著提升答案正确性;但提供多样文档极具益处,可将答案正确性提升17%-47%。我们进一步表明,这种提升仅由文档体裁(新闻、博客等)的多样性驱动,而非生成器可获取更相关答案的结果。我们的发现有助于引导更多关注新检索方法如何通过满足生成器对检索结果多样性的偏好来提升RAG。

英文摘要

In RAG, while the retriever typically ranks documents by their individual relevance to the query, the generator instead produces an answer based on the retrieved documents as a whole. This paper investigates how redundancy and diversity from the retrieved document set impact the generator in terms of answer correctness. Previous work has provided a mix of findings: some showing that redundancy improves generation by reinforcing relevant information, others that LLM-based paraphrasing of the same content may be beneficial. Many of these studies did not control for confounding factors like whether the documents contained the exact answer or not, and if parametric knowledge plays a role. We conduct a carefully controlled experiment investigating three key scenarios of retrieved document sets: 1) Duplicate (exact copies of the same document), 2) Paraphrased (LLM rephrased versions of one document) and 3) Diverse (documents from different genres each containing relevant information in different forms). We control for which documents contain the answer in exact match or rephrased form. Evaluation is done with FictionalQA, a synthetic, fictional question-answer dataset that ensures the LLM generator prior knowledge cannot answer the question; the answer must come from retrieved documents. We show that duplicate redundancy and LLM paraphrasing does not significantly improve answer correctness. However, providing diverse documents is highly beneficial, improving answer correctness by 17%-47%. We further show this improvement is driven by diverse forms of document genre (news, blogs, etc.) alone and not a consequence of more relevant answer being available to generator. Our findings help to direct more attention to how new retrieval methods might improve RAG by catering to the generator preference for diversity in retrieval results.

↑