arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

先获取再探索:面向搜索智能体的持久工作空间上选择与提取解耦

Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents

Qi Liu, Yiqun Chen, Zidan Chen, Yan Gao, Yi Wu, Yao Hu, Jiaxin Mao, Fengbin Zhu, Tat-Seng Chua

arXiv 2608.02097首次发表:更新:

发表机构

Renmin University of China; Xiaohongshu Inc.; National University of Singapore(中国人民大学; 小红书公司; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对搜索智能体提出Fetch-then-Explore方法,将页面选择与证据提取解耦,在两个基准测试中提升了智能体的搜索性能,核心是通过持久工作空间实现页面复用与证据累积。

AI 中文摘要

搜索智能体现在能够回答需要数十次搜索才能解决的问题,但相较于如何找到页面,智能体如何阅读页面的关注度要低得多。几乎所有这类智能体都使用两种文档界面中的一种,且两者都将页面与打开它的时刻绑定。“访问并阅读”在获取时将页面阅读内容注入消息历史,在智能体知道需要哪个事实之前就固定了该阅读内容;而有状态的“浏览”则按需从当前页面提取内容,但一次仅保留一个页面,且智能体打开另一个页面时就会释放该页面。无论哪种方式,许多轮后才发现重要的页面都必须重新获取并渲染到上下文中。我们提出“先获取再探索”(Fetch-then-Explore),它将页面选择与证据提取分离,并保留所选择的内容:页面记录在文件系统上针对每个问题的工作空间中,而非上下文窗口或临时会话中,证据可在之后按需从中提取。选择几乎无成本,提取可等到智能体明确要查找的内容后再进行,且可随假设细化重复,智能体切换时不会释放页面,因此证据会在整个轨迹中累积。在带有固定搜索的统一ReAct框架中,我们在两个开放网络基准BrowseComp和WideSearch上,针对三个智能体骨干,将Fetch-then-Explore与仅片段、访问并阅读、浏览基线进行比较。它在每个骨干上都领先BrowseComp准确率,在WideSearch上总体与基线相当或更优,行为分析将增益归因于工作空间的核心设计:离开页面后仍可返回,其返回页面的频率远高于任何临时界面,因此首次浏览时遗漏的证据仍可在之后恢复。

英文摘要

Search agents now answer questions that take dozens of searches to settle, yet how such an agent reads a page has drawn far less attention than how it finds one. Nearly all of them use one of two document interfaces, and both tie a page to the moment it is opened. \emph{Visit-and-read} injects a reading of the page into the message history at fetch time, fixing that reading before the agent knows which fact it will need. Stateful \emph{browsing} instead extracts on demand from the page in hand, but holds one page at a time and releases it as soon as the agent opens another. Either way, a page that turns out to matter many turns later has to be fetched and rendered into context all over again. We propose \textbf{Fetch-then-Explore}, which separates page selection from evidence extraction and keeps what it selects: pages are recorded in a per-question workspace on the filesystem rather than the context window or a transient session, and evidence is pulled from them on demand later. Selection becomes almost free, extraction can wait until the agent knows what to look for and be repeated as its hypothesis sharpens, and pages are not released when the agent moves on, so evidence accumulates across the trajectory. In a unified ReAct harness with fixed search, we compare Fetch-then-Explore against snippet-only, visit-and-read, and browsing baselines on two open-web benchmarks, BrowseComp and WideSearch, across three agent backbones. It leads BrowseComp accuracy at every backbone and generally matches or exceeds the baselines on WideSearch, and a behavioral analysis traces the gains to the workspace's defining move: returning to a page after leaving it, which it does far more than any transient interface, so evidence missed on a first pass can still be recovered later.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑