arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HarnessEvolve:利用参考轨迹实现智能体的可靠自进化

HarnessEvolve: Learning from Reference Trajectories for Reliable Agent Self-Evolution

Wen Jiang, Mingmin Chu, Yimeng Tian, Qianxin Zhang, Haofei Yang, Rui Yang, Yang Liu, Tao Lv, Fangming Li

arXiv 2609.00829首次发表:更新:

发表机构

ICT AI Competence Center, Huawei Technologies(华为技术有限公司ICT AI能力中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

HarnessEvolve是一种将执行与进化解耦的智能体自进化框架,通过参考轨迹克服信用分配失败,经质量与性能门控防止捷径学习和灾难性遗忘,在多基准上性能优于现有方法。

AI 中文摘要

自进化智能体通过基于环境反馈优化其控制组件(提示词、技能、工具及执行逻辑)向自主能力迈进。然而,该范式面临三大挑战:一是信用分配失败,即终端成功/失败反馈难以明确是哪一步导致了错误;二是捷径学习,即智能体记忆任务特定模式而非获取可泛化能力;三是灾难性遗忘,即无防护的更新会降低先前获得的能力。本文提出HarnessEvolve,一种利用参考轨迹实现智能体可靠自进化的框架。HarnessEvolve将执行智能体与进化流程解耦,将执行、评估、优化及门控分配给独立的智能体模块,实现可泛化且稳定的控制组件改进。具体而言,HarnessEvolve通过生成参考轨迹(给定标准答案时产生的执行路径),并将失败的执行与参考轨迹对齐以提取错误信号,对这些信号进行聚类以揭示系统性失败模式,从而克服信用分配失败。为防止捷径学习和灾难性遗忘,候选控制组件更新必须通过两个门:质量门,用于过滤数据泄露和提示词冗余;性能门,若更新在当前批次上有所改进且不会降低近期批次的性能则接受该更新,同时在保留验证集上进行周期结束验证以选择性能最佳的已接受智能体快照。我们在涵盖开放域和企业场景的多个基准上开展了广泛实验,使用不同模型和智能体框架。结果表明,HarnessEvolve在所有基准和设置下均始终优于最先进的基线,证实其在各任务领域的可靠性。

英文摘要

Self-evolving agents advance toward autonomy by optimizing their harness---prompts, skills, tools, and execution logic---based on environmental feedback. This paradigm, however, is hampered by three challenges: \textit{credit assignment failure}, where terminal success/failure feedback makes it ambiguous which step caused the error; \textit{shortcut learning}, where agents memorize task-specific patterns rather than acquire generalizable capabilities; and \textit{catastrophic forgetting}, where unguarded updates degrade previously acquired competence. In this paper, we introduce HarnessEvolve, a self-evolving framework that learns from reference trajectories to achieve reliable agent self-evolution. HarnessEvolve decouples the execution agent from the evolutionary pipeline, assigning execution, evaluation, optimization, and gating to independent agent modules, enabling generalizable and stable harness improvements. Specifically, HarnessEvolve overcomes credit assignment failure by generating reference trajectories (execution paths produced when given the ground-truth answers) and aligning failed executions against them to extract error signals, which are clustered to reveal systematic failure patterns. To prevent shortcut learning and catastrophic forgetting, candidate harness updates must pass two gates: a quality gate that filters data leakage and prompt bloat, and a performance gate that accepts each update if it improves on the current batch without degrading recent batches, with epoch-end validation on a held-out set selecting the best-performing accepted agent snapshot. We conduct extensive experiments on several benchmarks spanning open-domain and enterprise scenarios, using different models and agent frameworks. Results demonstrate that HarnessEvolve consistently outperforms state-of-the-art baselines across all benchmarks and settings, confirming reliability across task domains.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑