arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04968cs.LG

EvolveNet:智能体自我改进的协作式工具链进化

EvolveNet: Collaborative Harness Evolution for Agent Self-Improvement

  • Hong Kong Baptist University(香港浸会大学)
  • University of Science and Technology of China(中国科学技术大学)
  • The Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Jun Nie, Yonggang Zhang, Qianshu Cai, Yiu-ming Cheung, Xinmei Tian, Bo Han

AI总结:

EvolveNet提出协作式工具链进化范式,通过数据本地智能体的程序适配组合实现共享工具链改进,在五类任务场景均获增益,异构工作负载下效果更显著。

AI中文摘要:

大语言模型(LLM)智能体的能力不仅取决于其模型,还取决于工具链(harness):即构建上下文、调用工具、验证结果以及从故障中恢复的可执行程序。近期研究表明,进化工具链无需更新模型权重即可实现持续改进。然而,现有方法假设所有执行经验都可路由至单个优化器,该优化器沿单一轨迹进化一个工具链。但真实的智能体生态系统违背了这一假设:用户、组织和环境会生成无法汇集的孤立经验流,因此最值得学习的经验恰恰是无法直接集中的经验。我们提出EvolveNet,一种将经验提取转移至数据的协作式工具链进化范式。共享工具链被广播至数据本地的智能体部署节点,每个节点在自身工作负载上对其进行进化。仅将生成的程序适配部分组合为更新后的共享工具链并重新分发,使每个参与的智能体都能继承其他节点发现的操作经验。通过将聚合边界从原始工作负载转移至学习到的适配部分,EvolveNet保持工作负载本地化,允许多个进化搜索并发进行,同时降低串行深度。由于独立修改的程序无法像模型参数那样求平均,且组合时可能存在冲突,EvolveNet引入了范围类型、证据引导的程序聚合机制。在文本转SQL、数据科学编码、竞赛编程、软件工程和智能体工作流这五个场景中,EvolveNet均实现了共享工具链的性能提升,在异构工作负载下增益最大;消融实验表明,改进源于不同智能体的适配组合而非单纯的选择。

英文摘要:

The capabilities of an LLM agent depend not only on its model but on the harness: the executable program that constructs context, invokes tools, verifies results, and recovers from failure. Recent work shows that evolving the harness yields persistent improvements without updating model weights. Existing approaches, however, assume that all execution experience can be routed to a single optimizer, which evolves one harness along a sequential trajectory. Real agent ecosystems violate that assumption: users, organizations, and environments generate isolated streams of experience that cannot be pooled, so the experience most worth learning from is exactly the experience that cannot be directly centralized. We introduce EvolveNet, a paradigm of collaborative harness evolution that moves experience extraction to the data. A shared harness is broadcast to data-local agent deployments, each of which evolves it on its own workload. Only the resulting program adaptations are composed into an updated shared harness and redistributed, so that every participating agent inherits operational experience discovered by the others. By shifting the aggregation boundary from raw workloads to learned adaptations, EvolveNet keeps workloads local and allows multiple evolutionary searches to proceed concurrently with reduced serial depth. Because independently modified programs cannot be averaged like model parameters and may conflict when composed, EvolveNet introduces scope-typed, evidence-guided program aggregation. Across five settings spanning text-to-SQL, data-science coding, competitive programming, software engineering, and agentic workflows, EvolveNet improves the shared harness in all five, with the largest gains under heterogeneous workloads, and ablations attribute the improvement to composition of adaptations from different agents rather than to selecting among them.

补充信息

↑