SciRet:面向科学RAG的检索与重排序的计算感知实证研究
SciRet: A Compute-Aware Empirical Study of Retrieval and Reranking for Scientific RAG
浏览论文内容
中文总结 AI 辅助
该研究针对CORD-19科学问答开展SciRet实证研究,评估固定科学RAG流程在不同语料库规模下的表现,发现混合检索更稳健、MS MARCO训练的重排序器存在领域不匹配问题,生成忠实度随语料库规模提升,并发布相关资源支持后续研究。
中文摘要 AI 辅助
我们提出了SciRet,这是一项针对CORD-19科学问答的检索增强生成的计算感知实证研究。本文未提出新模型,而是在三个语料库规模下评估了固定的科学RAG流程:1034个块(1000篇论文)、5160个块(5000篇论文)和15480个块(15000篇论文)。该流程结合了句子窗口分块、BM25、BGE-M3密集检索、倒数排名融合、可选的交叉编码器重排序以及基于事实的答案生成。在我们的设置中,混合检索比仅稀疏或仅密集检索更稳健,在1000篇和15000篇论文规模下的Recall@10达到1.000。相比之下,在MS MARCO上训练的交叉编码器重排序器会降低科学语料库的精度,这表明领域不匹配的影响可能超过更强的查询-段落交互带来的益处。在我们的设置中,通过RAGAS衡量的生成忠实度随语料库规模增加而提升。检索评估使用来自混合系统的伪相关标签,因此我们将结果视为受控比较证据而非基准主张。我们发布代码、索引和评估输出,以支持可复现性和后续研究。
英文摘要
We introduce SciRet, a compute-aware empirical study of retrieval-augmented generation for scientific question answering over CORD-19. Rather than proposing a new model, we evaluate a fixed scientific RAG pipeline across three corpus scales: 1,034 chunks (1K papers), 5,160 chunks (5K papers), and 15,480 chunks (15K papers). The pipeline combines sentence-window chunking, BM25, BGE-M3 dense retrieval, reciprocal rank fusion, optional cross-encoder reranking, and grounded answer generation. Across these settings, hybrid retrieval is more robust than either sparse-only or dense-only retrieval in our setting, reaching Recall@10 of 1.000 at 1K and 15K. In contrast, an MS MARCO-trained cross-encoder reranker reduces precision on the scientific corpus, suggesting that domain mismatch can outweigh the benefits of stronger query-passage interaction. Generation faithfulness measured with RAGAS increases with corpus scale in our setup. Retrieval evaluation uses pseudo-relevance labels derived from the hybrid system, so we treat the results as controlled comparative evidence rather than a benchmark claim. We release code, indexes, and evaluation outputs to support replication and follow-up studies.