发表机构
IBM; IIT Kharagpur(国际商业机器公司; 克勒格布尔印度理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RIT-RAG结合内容检索与结构导航,通过构建检索诱导树解决智能体RAG的文档结构缺失问题,在多领域基准测试中准确率优于多种基线,在EntQABench上提升显著。
AI 中文摘要
检索增强生成(RAG)将语言模型与外部语料库绑定。智能体RAG支持迭代搜索,但会让模型接触到孤立的文本块,而这些文本块不包含文档结构,导致难以将相关证据与仅在表面上类似查询的文本块区分开。PageIndex等感知结构的方法可导航文档结构,但无法扩展到不适合LLM上下文的大型语料库结构,这类方法会先通过文档检索器确定单个文档,且无法从错误选择中恢复。因此,我们提出RIT-RAG(Retrieval-Induced Tree RAG),它将内容检索与结构导航相结合。在离线阶段,RIT-RAG会根据每个文档的目录或站点地图为其构建一棵树;在查询阶段,它会检索大量文本块,并利用这些文本块的位置来生成可管理的子树,这些子树可能跨多个文档。LLM智能体可导航这些子树,有选择地读取有前景的节点,并在需要时重新表述查询。由此,检索负责提出查找方向,而智能体决定读取内容。在金融、科学和客户支持基准测试中,RIT-RAG在普通基线、基于图的基线和智能体基线中实现了最高的答案准确率。在包含284万技术文档网页的新基准EntQABench上,它在三种LLM上的准确率比最强基线高出6.8至11.4个百分点。
英文摘要
Retrieval-augmented generation (RAG) grounds language models in external corpora. Agentic RAG enables iterative search, yet exposes the model to isolated chunks without document structure, making it difficult to distinguish relevant evidence from chunks that merely resemble the query. Structure-aware methods such as PageIndex navigate document structure but cannot scale to the structures of large corpora, which do not fit in the LLM context. Hence, they first commit to a single document using a document retriever and cannot recover from a wrong choice. We propose RIT-RAG (Retrieval-Induced Tree RAG), which combines content retrieval with structural navigation. Offline, RIT-RAG builds a tree for each document from its table of contents or sitemap. At query time, it retrieves a broad set of chunks and uses their positions to induce manageable sub-trees, potentially across multiple documents. An LLM agent navigates these sub-trees, selectively reads promising nodes, and reformulates queries when needed. Thus, retrieval proposes where to look, while the agent decides what to read. Across financial, scientific, and customer-support benchmarks, RIT-RAG achieves the highest answer accuracy among vanilla, graph-based, and agentic baselines. On EntQABench, our new benchmark of 2.84 million technical-documentation webpages, it improves accuracy by 6.8 to 11.4 points over the strongest baseline across three LLMs.
Comments24 pages (main text through Limitations ends on page 9, followed by references and appendix), 9 figures, 16 tables