arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35816cs.CLcs.AI

PrimeSeeker:面向深度搜索智能体的能力导向监督

PrimeSeeker: Capability-Oriented Supervision for Deep Search Agents

Linzhi Peng, Hanting Chen, Heng Chang, Ke Cheng, Bowen Du, Weifeng Lv

首次发表
浏览论文内容

中文总结 AI 辅助

针对深度搜索智能体训练中全局难度指标与局部检索能力不匹配的问题,提出PrimeSeeker框架,通过潜在锚点推理构建基于网络的锚点结构并联合生成问题与参考证据骨架,以监督微调和强化学习优化,在五个基准上以更少工具调用实现强性能。

中文摘要 AI 辅助

大型语言模型搜索智能体通常使用合成问题进行训练,这些问题的难度通过更大的证据图、更多的跳数和更长的轨迹来提升。然而,这些全局属性只是搜索过程中所需局部检索能力的间接代理。为解决这一不匹配问题,我们引入了潜在锚点推理,其内容包括从描述性规范中解析未命名的检索锚点,并将恢复的锚点转移到后续的信息需求中。这一原始检索单元将深度搜索分解为耦合操作的链,并围绕锚点解析和关系转移组织问题构建,而不规定规范搜索路径。基于此表述,我们提出了PrimeSeeker,一个面向能力的框架,该框架构建基于网络的锚点结构,并联合推导出一个问题和参考证据骨架。该骨架保留构建过程中的支持证据,并通过当前工具观测的提取性高亮来指导专家生成。这些高亮在监督微调前被移除,而骨架随后被重用,以审计参考步骤覆盖率,用于强化学习奖励。我们构建了9,221条专家轨迹,训练了一个30B搜索智能体。在五个深度搜索基准上,PrimeSeeker取得了强劲性能,而参考步骤优化进一步提升了监督策略。由此产生的轨迹表现出低检索冗余,固定预算评估显示,与长视界系统相比,其以显著更少的工具调用实现了强解决方案覆盖率。

英文摘要

Large language model search agents are often trained with synthetic questions whose difficulty is increased through larger evidence graphs, additional hops, and longer trajectories. These global properties, however, are only indirect proxies for the local retrieval capabilities required during search. To address this mismatch, we introduce latent anchor reasoning, which consists of resolving an unnamed retrieval anchor from descriptive specifications and transferring the recovered anchor into a subsequent information demand. This primitive retrieval unit decomposes deep search into chains of coupled operations and organizes question construction around anchor resolution and relation transfer, without prescribing a canonical search path. Based on this formulation, we propose PrimeSeeker, a capability-oriented framework that constructs web-grounded anchor structures and jointly derives a question and a reference evidence skeleton. The skeleton preserves supporting evidence from construction and guides expert generation through extractive highlights of current tool observations. These highlights are removed before supervised fine-tuning, while the skeleton is subsequently reused to audit reference-step coverage for reinforcement-learning rewards. We construct 9,221 expert trajectories, training a 30B search agent. Across five deep-search benchmarks, PrimeSeeker achieves strong performance, while reference-step optimization further improves the supervised policy. The resulting trajectories exhibit low retrieval redundancy, and fixed-budget evaluation shows strong solution coverage with substantially fewer tool calls than long-horizon systems.

发表机构

  • Beihang University(北京航空航天大学)
  • Huawei Technologies Ltd.(华为技术有限公司)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

↑