arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04889cs.AI

ForkPilot:面向长时程智能体的自演化回溯搜索策略

ForkPilot: Self-Evolving Policy for Retrospective Search in Long-Horizon Agents

Xinyue Zeng, Shivam Shandilya, Guilherme Potje, Leonardo Nunes, Rakshanda Agarwal, Ranveer Chandra, Emre Kiciman, Dawei Zhou, Tusher Chakraborty

首次发表
浏览论文内容

中文总结 AI 辅助

针对长时程智能体中回溯搜索的归因与适应复杂性,提出自演化两阶段策略框架ForkPilot,在6个基准和7个LLM家族上达到最先进性能并减少高达59.2%的令牌使用。

中文摘要 AI 辅助

交互式语言模型智能体日益通过长时程、多轮调用的推理来解决复杂任务,在此过程中,信念或行动中的错误可能随着工具交互而不断累积。回溯搜索能够从这类失败中恢复,但容易产生资源分配不当的问题。延迟的结果掩盖了中间搜索决策的贡献,导致归因复杂性;而不断演化的执行证据则导致适应复杂性,使先前学习的估计变得过时。为应对这些挑战,我们首先引入搜索价值动态(SVD),用以刻画回溯搜索的收益与成本之间不断演化的权衡。基于SVD,我们提出ForkPilot,一个自演化的两阶段策略学习框架。在第一阶段,ForkPilot通过自动构建的结果比较,从已完成的轨迹中离线学习搜索价值策略。在第二阶段,它基于当前观测做出搜索决策,然后通过将新完成的轨迹纳入后续策略更新中实现自我演化。我们在6个多样化的基准测试和7个广泛使用的LLM骨干家族(包括四个开源家族、GPT-5.6 Sol和Opus 4.8)中,在生产级智能体系统中,与9个竞争性基线(包括一个被数十万付费用户使用的真实世界工具部署)进行了评估。ForkPilot在达到可比的最先进性能的同时,将令牌使用量减少了高达59.2%,证明了其有效性。

英文摘要

Interactive language-model agents increasingly solve complex tasks through long-horizon, multi-call reasoning, where errors in beliefs or actions can compound across tool interactions. Retrospective search can recover from such failures but is prone to misallocation. Delayed outcomes obscure the contribution of intermediate search decisions, leading to Attribution Complexity, while evolving execution evidence leads to Adaptation Complexity, where previously learned estimates become stale. To address these challenges, we first introduce Search Value Dynamics (SVD), which characterizes the evolving trade-off between the gain and cost of retrospective search. Building on SVD, we propose ForkPilot, a self-evolving two-stage policy-learning framework. In the first stage, ForkPilot learns a search-value policy offline from completed trajectories through automatically constructed outcome comparisons. In the second stage, it makes search decisions based on current observations and then self-evolves by incorporating newly completed trajectories into subsequent policy updates. We evaluate ForkPilot across 6 diverse benchmarks and 7 widely used LLM backbone families, including four open-source families, GPT-5.6 Sol, and Opus 4.8 in a production agentic system, against 9 competitive baselines, including a real-world harness deployment used by hundreds of thousands of paid users. ForkPilot achieves comparable state-of-the-art performance while reducing token usage by up to 59.2%, demonstrating its efficacy.

发表机构

  • Virginia Tech(弗吉尼亚理工大学)
  • Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

↑