arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DREvo:提炼重校准的历史经验以实现工具链自主进化

DREvo: Distilling Recalibrated Historical Experience for Harness Self-Evolution

Hanghui Guo, Weijie Shi, Zhangze Chen, Shengxiang Xu, Yishu Wang, Yimei Zhang, Wangze Ni, Jia Zhu, Shimin Di

arXiv 2607.26722首次发表:更新:

发表机构

Southeast University; Hong Kong University of Science and Technology; Zhejiang Normal University; Zhejiang University of Technology; Zhejiang University(东南大学; 香港科技大学; 浙江师范大学; 浙江工业大学; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有工具链自主进化方法的两个局限,提出DREvo方法,通过整合三项技术提升进化稳定性,在五个基准上取得最优准确率,在两类任务上较基线实现显著性能提升。

AI 中文摘要

工具链(Harness)在大语言模型智能体的性能中起着关键作用,构建高性能工具链需要大量专家精力。因此,近期研究越来越多地探索工具链自主进化,即利用历史试错经验迭代地提出、评估和改进工具链。然而,积累的历史经验并不总能转化为稳定的搜索指导,且在进化迭代中性能往往大幅波动,导致在有限的进化预算下难以可靠地发现高性能工具链。我们发现现有工具链自主进化方法在利用历史经验时存在两个局限:(1)缺乏对历史经验是否仍适用于当前工具链的动态重新评估;(2)缺乏将有效历史经验转化为可操作搜索方向的明确机制。为解决这些局限,我们提出一种名为DREvo的新工具链自主进化方法,它整合了功能级证据锚定、状态依赖的证据重校准以及角色条件的搜索意图提炼,以确定哪些历史证据仍然有效以及工具链下一步应向何处进化。在有限的进化预算下,DREvo呈现出更平稳的进化轨迹,在所有五个基准上达到最高准确率,且在领域推理和智能体任务上分别比评估的基线方法实现了16.2%和14.2%的平均提升。

英文摘要

Harness plays a critical role in large language model agent performance, and building a high-performing harness requires substantial expert effort. Therefore, recent research has increasingly explored harness self-evolution, which iteratively proposes, evaluates, and improves harnesses using historical trial experience. However, accumulated historical experience does not always translate into stable search guidance, and performance often fluctuates substantially across evolution iterations, making it difficult to reliably discover high-performing harnesses under a limited evolution budget. We identify two limitations in how existing harness self-evolution methods leverage historical experience: (1) Lack of dynamic reassessment of whether historical experience remains valid for the current harness, and (2) Lack of explicit mechanisms for translating valid historical experience into actionable search directions. To address these limitations, we propose a new harness self-evolution method, named DREvo, which integrates function-level evidence anchoring, state-dependent evidence recalibration, and role-conditioned search intent distillation to determine which historical evidence remains valid and where the harness should evolve next. Under limited evolution budgets, DREvo exhibits smoother evolution trajectories, achieves the highest accuracy on all five benchmarks, and delivers average gains of 16.2% and 14.2% over the evaluated baselines on domain reasoning and agentic tasks, respectively.

Comments9 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑