发表机构
Xi’an Jiaotong-Liverpool University; Suzhou Lab; China University of Petroleum (Beijing); University of Liverpool(西交利物浦大学; 苏州实验室; 中国石油大学(北京); 利物浦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对科学问答中识别相关论文及找支持证据的问题,提出VecTree-RAG框架,结合向量与树检索机制。经多组问题评估,该框架在多个基准上获高分,证据页面精度高,且完整架构所需推理令牌少,为科学文献问答提供结构感知且可追溯的架构。
AI 中文摘要
科学问答需要一个检索系统来解决两个不同的问题:识别哪些论文相关,并在这些论文中找到支持证据。传统的检索增强生成通常通过对固定长度段落进行相似性搜索来解决这两个问题,扁平化文档结构并将科学主张与其方法和论证背景分开。我们提出了VecTree-RAG,一个将这些任务分配给互补检索机制的智能框架。向量搜索对语料库中的紧凑文档和部分表示进行排名,而对源验证部分树的推理引导遍历则在入围论文中定位证据。全文保留在页面存储中,仅在结构定位后逐步显示。我们在300个QASPER问题、54个LitQA2问题的开放访问子集和49个多文档MOSAIC问题上评估了VecTree-RAG。与Dense RAG、重新排序的Dense RAG、RAPTOR和Search-o1相比,VecTree-RAG在所有三个基准上都获得了最高的观察答案分数,在QASPER上达到了0.800的LLM判断正确性,在LitQA2上达到了0.925的准确率,在MOSAIC上达到了0.547的综合分数。在QASPER上,其证据页面精度为0.274,而基线为0.046-0.071。LitQA2的消融实验进一步表明,完整的向量-树架构比没有树导航或语料库级向量路由的变体需要更少的推理令牌。这些结果表明,向量检索缩小了语料库级搜索空间,树导航将阅读集中在结构相关的证据上。虽然多轮推理仍然比单次调用检索更昂贵,但VecTree-RAG为科学文献问答提供了一种结构感知和可追溯的架构。
英文摘要
Scientific question answering requires a retrieval system to solve two distinct problems: identifying which papers are relevant and locating the supporting evidence within those papers. Conventional retrieval-augmented generation typically addresses both through similarity search over fixed-length passages, flattening document structure and separating scientific claims from their methodological and argumentative context. We present VecTree-RAG, an agentic framework that assigns these tasks to complementary retrieval mechanisms. Vector search ranks compact document and section representations across the corpus, whereas reasoning-guided traversal of source-verified section trees localizes evidence within shortlisted papers. Full text is retained in a page store and exposed progressively only after structural localization. We evaluate VecTree-RAG on 300 QASPER questions, an open-access subset of 54 LitQA2 questions, and 49 multi-document MOSAIC questions. Compared with Dense RAG, reranked Dense RAG, RAPTOR, and Search-o1, VecTree-RAG obtained the highest observed answer score on all three benchmarks, reaching 0.800 LLM-judge correctness on QASPER, 0.925 accuracy on LitQA2, and a 0.547 composite score on MOSAIC. On QASPER, its evidence-page precision was 0.274, compared with 0.046--0.071 for the baselines. LitQA2 ablations further showed that the complete vector--tree architecture required fewer inference tokens than variants without tree navigation or corpus-level vector routing. These results indicate that vector retrieval narrows the corpus-level search space and tree navigation concentrates reading on structurally relevant evidence. Although multi-turn inference remains more expensive than single-call retrieval, VecTree-RAG provides a structure-aware and traceable architecture for scientific literature question answering.