直接语料库交互中的证据失明:基于AtlasNav的持久导航
AtlasNav: Mitigating Evidence Blindness with Persistent Corpus Navigation
AI总结:
针对直接语料库交互中存在的证据失明问题,提出AtlasNav框架,通过一次性构建语料库图谱实现自适应导航,在BrowseComp-Plus等数据集上提升了准确率并降低了推理成本。
AI中文摘要:
大型语言模型智能体正从传统的检索增强生成转向与外部语料库的直接交互(Direct Corpus Interaction, DCI)。DCI可让完整语料库保持可访问状态,但在有限的交互预算下,可获取的证据可能仍无法使用:所需证据可能未浮现,浮现的支持文档可能未被打开,或已打开的文档可能无法暴露其决定性片段。我们将这种渐进式的无声损失称为“证据失明”,并通过分阶段的证据实现对其进行量化。在DCI范式中,原始交互几乎不会产生可复用的语料库组织,而动态工作空间方法会针对每个查询和轨迹重构以查询为条件的交互空间,在这两种情况下,有用的结构大多是在线恢复的。我们则将大规模智能体搜索建模为在可复用语料库结构上的有限预算导航,提出AtlasNav——一种持久的多视图语料库导航框架,它保留了直接语料库交互,但会将语料库一次性组织为语料库图谱(Corpus Atlas),使每个查询能够自适应导航,而非重构共享结构。在BrowseComp-Plus上,AtlasNav实现了92.05%的严格准确率,相较于之前的动态工作空间SOTA,可降低30.21%的记录在线推理成本;在匹配的预算下,它能更早实现完整所需证据,并更快接近同一模型的证据提供的经验参考。该相同表示原则在具有不同语料库组织的PhantomWiki及受控的10K-1M规模下仍有效,且可竞争性地迁移到异构企业知识中。这些结果表明,智能体搜索不仅依赖于可访问的证据,还依赖于语料库的表示方式,以让有限的交互成为有效的导航。
英文摘要:
As language-model agents become more capable of iterative search, corpus access is shifting from retrieval toward interaction. Agents can explore the corpus, inspect documents, and use newly discovered evidence to decide what to examine next. Yet accessible evidence may still fail to become usable within a finite interaction budget. We call this progressive failure Evidence Blindness: supporting documents may never enter view, may remain unopened, or may fail to expose the decisive evidence even after being opened. A key reason is that agents often have to infer useful evidence directions during interaction, spending limited budget on deciding where to search next. Existing approaches either leave corpus structure largely implicit or reconstruct useful directions at query time. AtlasNav instead organizes reusable cross-document structure before any query arrives. It builds a persistent multi-view Corpus Atlas, which each query can navigate adaptively while still accessing the original documents directly. On BrowseComp-Plus, AtlasNav outperforms the previous state-of-the-art interactive corpus access method across different backbones. On DeepSeek, it improves strict accuracy by 7.47 points while reducing query-time inference cost by 30.22%.AtlasNav also reduces Evidence Blindness, realizes complete evidence earlier, remains robust to corpus-structure and scale shifts on PhantomWiki, and achieves leading performance on heterogeneous enterprise data.