arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Traverse:学习何时记忆、重置与重定向以进行长时程网络搜索

Traverse: Learning When to Remember, Reset, and Redirect for Long-Horizon Web Search

Jingyuan Ma, Lynx Aster, He Zhang, Siyao Song, Weijie Yuan, Zhe Zhang, Kai Jia, Zhifang Sui

arXiv 2609.37082首次发表:更新:

发表机构

Peking University; ByteDance(北京大学; 字节跳动)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长时程搜索中噪声累积问题,提出三状态自主搜索框架与Seal Memory工具,通过分段强化学习训练,35B模型在BrowseComp等基准上取得领先性能。

AI 中文摘要

长时程信息寻求智能体常常积累噪声或误导性上下文,导致早期错误持续存在,并使恢复变得越来越困难。我们引入了一个自主搜索框架,其中智能体通过三种状态管理其自身的搜索过程:Rubric(标准)、Answer(回答)和Verify(验证)。智能体首先定义有效答案的标准,在这些标准下进行搜索,然后在决定终止或继续搜索之前独立验证结果。它进一步配备了Seal Memory(密封记忆)工具,以实现主动的上下文管理。然而,使用强化学习训练这种行为可能引发Seal Collapse(密封崩溃),导致训练不稳定,并阻止智能体可靠地学习何时以及如何使用其记忆工具。我们通过一个简单的策略解决了这一问题,即仅在上下文管理之后训练最终阶段。我们的35B模型在BrowseComp上达到72.83分,优于可比较的开源系统,并在BrowseComp-ZH、xbench、DeepSearchQA、WideSearch、金融调查和产品搜索上持续优于基础模型。消融实验表明,自主压缩优于自动压缩,并验证了我们的强化学习设计。

英文摘要

Long-horizon information-seeking agents often accumulate noisy or misleading context, causing early mistakes to persist and making recovery increasingly difficult. We introduce an autonomous search harness in which the agent manages its own search process through three states: Rubric, Answer, and Verify. The agent first defines criteria for a valid answer, searches under these criteria, and then independently verifies the result before deciding whether to terminate or continue searching. It is further equipped with a Seal Memory tool that enables active context management. Training this behavior with reinforcement learning, however, can induce Seal Collapse, resulting in unstable training and preventing the agent from reliably learning when and how to use its memory tools. We solve this with a simple strategy that trains only the final segment after context management. Our 35B model achieves 72.83 on BrowseComp, outperforming comparable open-source systems, and consistently improves over the base model across BrowseComp-ZH, xbench, DeepSearchQA, WideSearch, financial investigation, and product search. Ablations show that autonomous compression outperforms automatic compaction and validate our RL design.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑