发表机构
Renmin University of China; National University of Singapore(中国人民大学; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过轨迹级诊断区分长视野搜索智能体的检索与利用差距,发现搜索努力与答案质量弱相关,为优化深度研究系统提供了查询构建、证据管理等方向。
AI 中文摘要
深度搜索智能体通过迭代发出搜索查询以收集支撑证据,来回答困难的信息检索问题,但目前尚不清楚更多的搜索努力是否以及如何带来更好的答案。我们通过对长视野搜索智能体的轨迹级诊断来研究这些问题。利用人工标注的文档级相关性判断,我们评估每个搜索步骤检索到的证据,并将智能体行为分为两个阶段:智能体检索到什么证据,以及它如何有效利用这些证据。这种区分进一步使我们能够将失败分解为检索差距(即从未找到必要证据)和利用差距(即已检索到相关证据但未正确使用)。在检索模型和评估工具固定的情况下,我们在BrowseComp-Plus上比较了六个智能体,并使用开放网络搜索API在BrowseComp上验证了我们的发现。在所有设置中,我们发现搜索努力和答案质量仅呈弱相关。答案准确性与检索证据的质量(尤其是累积检索召回率)的相关性,比与搜索次数或消耗的上下文量的相关性更强。有用的证据常出现在轨迹早期,但智能体往往会继续搜索,产生低收益检索步骤的长尾。在查询层面,探索性的查询重述仍有用,但表现最佳的智能体发出的冗余查询少得多。总体而言,通过系统表征长视野搜索智能体的搜索行为和失败模式,本工作为构建更好的深度研究系统指明了实用方向,包括更强的查询构建、更有效的证据选择与上下文管理,以及基于是否已检索到足够支撑证据的停止准则。
英文摘要
Deep search agents answer difficult information-seeking questions by iteratively issuing search queries to gather supporting evidence, but it remains unclear whether and how greater search effort leads to better answers. We study these questions through a trajectory-level diagnosis of long-horizon search agents. Using human-annotated document-level relevance judgments, we evaluate the evidence retrieved at each search step and separate two stages of agent behavior: what evidence an agent retrieves and how effectively it uses that evidence. This distinction further allows us to decompose failures into retrieval gaps, where the necessary evidence is never found, and utilization gaps, where relevant evidence is retrieved but not used correctly. With the retrieval model and evaluation harness held fixed, we compare six agents on BrowseComp-Plus and further validate our findings on BrowseComp with an open-web search API. Across settings, we find that search effort and answer quality are only weakly aligned. Answer accuracy is better correlated with the quality of retrieved evidence, especially cumulative retrieval recall, than with the number of searches or the amount of context consumed. Useful evidence often appears early in the trajectory, yet agents tend to continue searching, producing a long tail of low-yield retrieval steps. At the query level, exploratory reformulations remain useful, but the best-performing agents issue far fewer redundant queries. Overall, by systematically characterizing the search behavior and failure modes of long-horizon search agents, this work points to practical directions for building better deep research systems, including stronger query formulation, more effective evidence selection and context management, and stopping criteria based on whether sufficient supporting evidence has been retrieved.