arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38473cs.IRcs.AIcs.LG

重排序与后期交互驱动检索质量:面向科学问答的RAG策略受控比较

Re-ranking and Late Interaction Drive Retrieval Quality: A Controlled Comparison of RAG Strategies for Scientific Question Answering

  • San Jose State University(圣何塞州立大学)

机构由 AI 辅助整理,请以论文原文为准。

Bhagyesh Rathi, Eshan Chawla, William B. Andreopoulos

中文总结 AI 辅助

本研究受控比较六种RAG检索策略,发现重排序与后期交互显著提升科学问答检索质量,并发布含19,484个合成问题的开放测试平台。

中文摘要 AI 辅助

检索增强生成(RAG)现已成为将大型语言模型(LLM)锚定于外部知识的标准方式,然而检索流程的设计空间庞大,各变体之间的权衡在现实规模的专业语料库上尚未得到充分理解。在本工作中,我们针对科学问答提出了六种检索策略的受控比较:(i)经典top-k稠密检索,(ii)基于LLM的查询改写,(iii)查询改写后接基于LLM的重排序,(iv)通过倒数排名融合(RRF)进行多查询融合,(v)一种智能体工具调用流水线,其中生成器自行决定是否检索,以及(vi)使用ColBERTv2的后期交互检索。所有六种流水线共享相同的生成器(Meta-Llama/Llama-3.1-8B-Instruct)、提示词和评估协议;五种单向量流水线额外共享SPECTER2嵌入和Chroma向量存储;所有六种均从包含463,971篇2024-2025年arXiv论文的完整语料库中检索。为支持可复现的大规模评估,我们还发布了一个合成问题数据集,包含由Llama-3.1-8B-Instruct从跨学术领域的10,000篇论文随机样本中生成的19,484个问题陈述和方法学问题(其中9,742个查询生成成功),每种策略均在此同一查询集上评估。我们描述了每种流水线的架构和实现,发布了代码和合成问题数据集,并使用LLM作为评判者的协议从多个质量维度评估每种策略,同时结合直接的金标准论文检索指标。其结果是为研究文献语料库上RAG设计选择的成本和质量权衡提供了一个开放测试平台,并为未来关于忠实性、检索鲁棒性和智能体检索的研究奠定了基础。

英文摘要

Retrieval-Augmented Generation (RAG) is now the standard way to ground Large Language Models (LLMs) in external knowledge, yet the design space of retrieval pipelines is large and the trade-offs between variants are not well understood, especially on domain-specific corpora at realistic scale. In this work, we present a controlled comparison of six retrieval strategies for scientific question answering: (i) classic top-k dense retrieval, (ii) LLM-based query rephrasing, (iii) query rephrasing followed by LLM-based reranking, (iv) multi-query fusion via Reciprocal Rank Fusion (RRF), (v) an agentic tool-call pipeline in which the generator decides for itself whether to retrieve, and (vi) late-interaction retrieval with ColBERTv2. All six pipelines share the same generator (Meta-Llama/Llama-3.1-8B-Instruct), prompt, and evaluation protocol; the five single-vector pipelines additionally share SPECTER2 embeddings and a Chroma vector store; and all six retrieve from the full corpus of 463,971 arXiv papers dated 2024-2025. To support reproducible, large-scale evaluation, we also release a synthetic question dataset of 19,484 problem-statement and methodology questions generated by Llama-3.1-8B-Instruct from a random sample of 10,000 papers across academic domains (query generation succeeded for 9,742 of them), and every strategy is evaluated on this same query set. We describe the architecture and implementation of each pipeline, release the code and the synthetic question dataset, and evaluate each strategy with an LLM-as-a-judge protocol along multiple quality dimensions, together with direct gold-paper retrieval metrics. The result is an open testbed for studying the cost and quality trade-offs of RAG design choices on a research-literature corpus, and a basis for future work on faithfulness, retrieval robustness, and agentic retrieval.

补充信息

↑