arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34988cs.CL

在正确的步骤提供正确的经验:为自进化智能体推导控制更新

The Right Lesson at the Right Step: Deriving Control Updates for Self-Evolving Agents

  • Sun Yat-sen University(中山大学)
  • Tsinghua University(清华大学)
  • Nankai University(南开大学)
  • University of Pennsylvania(宾夕法尼亚大学)
  • Peng Cheng Laboratory(鹏城实验室)

机构由 AI 辅助整理,请以论文原文为准。

Yunhe Su, ZiYi Dong, Tong Yu, Weijian Deng, Hao Li, Bowen Jiang, Pengxu Wei

AI总结:

针对自进化智能体经验重用缺乏局部控制的问题,提出EvoCUE框架,将智能体表示为状态机控制器,从执行轨迹学习控制更新,在精确位置插入经验,显著提升长工具使用任务的完成质量。

AI中文摘要:

自进化智能体通过重用过去的经验(通常以全局提示、记忆或反思的形式)来改进未来的行为。然而,这些机制很少控制经验生效的位置。在长时间的工具使用工作流中,相同的经验教训可能纠正一个决策,却干扰另一个决策,这使得经验重用成为一个局部控制问题,而不仅仅是记忆问题。我们提出了EvoCUE(基于证据的控制更新进化框架),这是一个从完成的智能体执行中学习可重用控制程序更新的框架。EvoCUE将智能体表示为一个显式的状态机控制器,其节点执行模型或工具调用,其边定义控制流下一步的传递位置。这使得工作流可在精确位置进行编辑,因此每个学习到的更新可以指定添加什么、在何处生效以及何时应用。从完成的轨迹中,EvoCUE使用残差目标和观察到的执行轨迹来提出局部指令或技能编辑。每个候选方案在其将生效的位置进行评估,方法是从相同的检查点恢复父控制器和编辑后的控制器,并比较它们的最终结果。被接受的编辑与适用规则一起编译,在留出的任务上确认,并由后续执行继承。我们在长时间的工具使用环境中评估EvoCUE,在这些环境中,学习到的约定必须到达正确的执行步骤。从一个没有基准特定入门指令的最小AppWorld控制器出发,EvoCUE学习了缺失的任务完成约定,并在Test-Normal和Test-Challenge上显著提高了成功率。在PAST-Bench办公工作流中,EvoCUE将组织要求从先前的情节转移到后续任务,提高了任务执行质量。这些结果表明,自进化智能体应将经验置于控制流中,而不是仅将其存储为文本。

英文摘要:

Self-evolving agents improve future behavior by reusing past experience, typically as global prompts, memories, or reflections. Yet these mechanisms rarely control where experience takes effect. In long tool-use workflows, the same lesson may correct one decision but distract another, making experience reuse a problem of localized control rather than memory alone. We introduce EvoCUE (Evolution through Control Updates from Evidence), a framework for learning reusable control-program updates from completed agent executions. EvoCUE represents the agent as an explicit state-machine controller, whose nodes perform model or tool calls and whose edges define where control passes next. This makes the workflow editable at precise locations, so each learned update can specify what to add, where it acts, and when it applies. From completed trajectories, EvoCUE uses residual goals and observed execution traces to propose localized instruction or skill edits. Each candidate is evaluated at the point where it would act by resuming the parent and edited controllers from the same checkpoint and comparing their final outcomes. Accepted edits are compiled with applicability rules, confirmed on held-out tasks, and inherited by later executions. We evaluate EvoCUE on long tool-use environments where learned conventions must reach the right execution step. From a minimal AppWorld controller without benchmark-specific onboarding instructions, EvoCUE learns the missing task-completion convention and substantially improves success on Test-Normal and Test-Challenge. On PAST-Bench office workflows, EvoCUE transfers organizational requirements from prior episodes to later tasks, improving task-execution quality. These results show that self-evolving agents should place experience inside the control flow, rather than only store it as text.

补充信息

↑