跟随实体:面向智能体检索的语料库地图
Follow the Entities: A Corpus Map for Agentic Search
- KAIST(韩国科学技术院)
- Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对大型文档集合中证据分散的问题,提出CorpusMap导航层,以实体为中心组织语料库,构建实体-文档图,使智能体高效收集证据,提升检索质量并减少token消耗。
AI中文摘要:
在大型文档集合上回答问题并完成任务,通常需要连接分散在多个文档中的证据,例如一个项目的批准记录在一个文档中,其需求在另一个文档中,而其最新状态在第三个文档中。最近的LLM智能体通过迭代搜索整个语料库来解决这一问题,而不是仅阅读一组固定的排名靠前的文档。然而,当语料库仅以平面文件集合的形式呈现时,一个相关文档无法表明它与其他文档的关系,因此智能体必须为每个查询重新发现这些关系,常常遗漏互补证据,同时消耗大量额外的token。为解决这一问题,我们引入了CorpusMap,一个围绕其重复出现的实体组织语料库的导航层,这些实体可从文档本身识别,并能将单个文档链接到跨来源的许多其他文档。具体而言,CorpusMap将每个重复出现的实体表示为一个实体页面,该页面聚合关于该实体的信息,并链接到引用该实体的每个文档,从而在实体和文档之间形成一个图,智能体可以遍历该图以收集原本分散的证据。此外,由于CorpusMap是通过离线解析跨文档中同一实体的提及而构建的,其链接在查询之间共享,而不是在推理时反复重新发现。使用7种不同模型和3个基准数据集,我们表明CorpusMap在证据发现和答案质量上均优于原始语料库的智能体搜索,同时平均使用更少的token,并进一步优于4种替代导航层,表明实体可作为导航大型文档集合的有效锚点。
英文摘要:
Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its latest status in a third. Recent LLM agents approach this by iteratively searching the full corpus rather than reading only a fixed set of top-ranked documents. However, when the corpus is exposed only as a flat collection of files, a relevant document gives no indication of how it relates to others, so the agent must rediscover these relationships for every query, often missing complementary evidence while simultaneously consuming substantial additional tokens. To address this, we introduce CorpusMap, a navigation layer that organizes the corpus around its recurring entities, which are identifiable from the documents themselves and can link a single document to many others across sources. Specifically, CorpusMap represents each recurring entity as an Entity Page that aggregates information about it and links to every document that refers to it, forming a graph between entities and documents that the agent can traverse to gather otherwise disconnected evidence. Moreover, since CorpusMap is constructed offline by resolving mentions of the same entity across documents, its links are shared across queries rather than rediscovered repeatedly at inference time. Using 7 different models with 3 benchmark datasets, we show that CorpusMap improves both evidence discovery and answer quality over raw-corpus agentic search while using fewer tokens on average, and further outperforms 4 alternative navigation layers, suggesting that entities serve as effective anchors for navigating large document collections.