发表机构
City University of Hong Kong; Meituan; University of Oxford(香港城市大学; 美团; 牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对深度研究智能体易因单一搜索轨迹导致失败的问题,提出HypoSearch算法,通过假设引导的多分支搜索与证据比较提升性能,在多基准和模型上优于基线,且工具调用更少。
AI 中文摘要
深度研究智能体通过与搜索和浏览工具交互来回答复杂问题,但它们通常沿着单一的演化轨迹进行搜索。我们对轨迹层面的分析发现了一种常见的失败模式:智能体可能在早期搜索状态下遇到多个可行方向,但在收集足够的比较证据之前就选择了其中一个方向。一旦发生这种情况,后续的工具调用往往会强化同一轨迹,若初始方向存在误导,失败的概率就会增加。我们进一步发现,成功的轨迹通过两种行为降低这种风险:将模糊的探索建立在具体候选的基础上,以及当当前路径薄弱或不完整时转换方向。基于这些发现,我们提出了HypoSearch,它生成轻量级假设作为软搜索提示,通过有界独立分支探索这些假设,并在决策前比较分支层面的证据。在四个深度研究基准和三个骨干模型上,HypoSearch始终优于单轨迹搜索和标准并行基线,在BC-small上将Qwen3.5-122B的性能从46.7提升至60.0,同时比五个独立轨迹使用更少的工具调用。一项初步的监督微调研究进一步表明,这些行为信号可以整理出紧凑的训练轨迹,并减少未过滤数据带来的性能下降。
英文摘要
Deep-research agents answer complex questions by interacting with search and browsing tools, yet they often search along a single evolving trajectory. Our trajectory-level analysis reveals a common failure mode in which the agent may encounter an early search state with several plausible directions, but follow one direction before collecting enough comparative evidence. Once this happens, subsequent tool calls tend to reinforce the same path, increasing the chance of failure when the initial direction is misleading. We further find that successful trajectories reduce this risk through two behaviors: grounding vague exploration in concrete candidates and shifting directions when the current path is weak or incomplete. Based on these findings, we propose HypoSearch, which generates lightweight hypotheses as soft search hints, explores them through bounded independent branches, and compares branch-level evidence before commitment. Across four deep-research benchmarks and three backbone models, HypoSearch consistently outperforms single-trajectory search and standard parallel baselines, improving Qwen3.5-122B from 46.7 to 60.0 on BC-small while using fewer tool calls than five independent trajectories. A pilot supervised fine-tuning study further shows that these behavioral signals can curate compact training trajectories and reduce degradation from unfiltered data.