程序化搜索智能体:将智能体搜索扩展到查询改写之外
Programmatic Search Agents: Extending Agentic Search Beyond Query Reformulation
浏览论文内容
中文总结 AI 辅助
针对搜索智能体无法控制候选处理与证据展示的问题,提出程序化搜索智能体(PSA),以本地可执行计算为搜索动作单元,统一候选工作区、灵活操作组合与选择性证据展示,在InfoSeek-Eval和BrowseComp-Plus上分别提升任务成功率4.00和7.56个百分点。
中文摘要 AI 辅助
搜索智能体会调整其查询,但固定的搜索界面使得候选处理与证据展示不在智能体的直接控制范围内。我们的轨迹分析表明,支持性段落可以被检索到却从未传递给智能体;同页面的预言机干预显示,改变返回的证据可以减少后续搜索。我们引入了程序化搜索智能体(PSA),它将候选上的本地可执行计算作为搜索动作的基本单元。PSA统一了持久候选工作区、灵活的原始操作组合和选择性证据展示。它增量生成程序单元,这些单元重用候选、执行依赖操作,并选择智能体接下来检查的内容。运行时在每个单元内解析指定的数据依赖,而智能体则随着新证据的到来在各单元间调整其搜索策略。我们在InfoSeek-Eval和BrowseComp-Plus上使用五种策略骨干(无需任务特定训练)将PSA与基于查询的智能体和基于工具的智能体进行比较。所有三种界面共享搜索底层,基于工具的智能体还共享PSA的原始操作和持久工作区。相对于基于查询的智能体,PSA在两个基准上的宏平均任务成功率分别提高了4.00和7.56个百分点;骨干内的最终步骤令牌平均减少了28.3%和33.9%。这些结果支持将智能体控制扩展到查询改写之外,涵盖检索证据的处理与展示。代码将在批准后发布。
英文摘要
Search agents adapt their queries, yet fixed search interfaces leave candidate processing and evidence presentation outside the agent's direct control. Our trajectory analysis shows that supporting passages can be retrieved yet never delivered to the agent; a same-page oracle intervention shows that changing the returned evidence can reduce subsequent search. We introduce Programmatic Search Agent (PSA), which makes a local executable computation over candidates the unit of a search action. PSA unifies a persistent candidate workspace, flexible primitive composition, and selective evidence presentation. It incrementally generates program cells that reuse candidates, execute dependent operations, and select what the agent inspects next. The runtime resolves specified data dependencies within each cell, while the agent adapts its search strategy across cells as new evidence arrives. We compare PSA with the Query-based Agent and Tool-based Agent on InfoSeek-Eval and BrowseComp-Plus using five policy backbones without task-specific training. All three interfaces share the search substrate, and the Tool-based Agent also shares PSA's primitives and persistent workspace. Relative to the Query-based Agent, PSA improves macro-averaged task success by 4.00 and 7.56 percentage points on the two benchmarks, respectively; within-backbone reductions in final-step tokens average 28.3% and 33.9%. These results support extending agent control beyond query reformulation to the processing and presentation of retrieved evidence. Code will be released subject to approval.
发表机构
- Zhejiang University(浙江大学)
- Yuanbao Team, Tencent(腾讯元宝团队)
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。