arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

反思自进化智能体技能:多轮次的反馈动态

Rethinking Self-Evolving Agent Skills: Feedback Dynamics over Multiple Rounds

Yuxuan Liu, Zhaochen Su, Yuhao Zhang, Jiahe Guo, Zhongwei Xie, Huihao Jing, Lingyun Xie, Qing Zong, Yauwai Yim, Zhixiong Zhang, Haoran Li, Yangqiu Song

arXiv 2608.02636首次发表:更新:

发表机构

The Hong Kong University of Science and Technology; Harbin Institute of Technology; Harbin Institute of Technology, Shenzhen; Shanghai Jiao Tong University(香港科技大学; 哈尔滨工业大学; 哈尔滨工业大学(深圳); 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出受控评估框架,发现自进化智能体技能是稀疏的验证过滤搜索,收益依赖模型与基准,失败轨迹反馈对技能选择关键,测试时计算难以完全恢复其收益。

AI 中文摘要

自进化技能系统有望通过将执行反馈转化为持久的技能更新来改进智能体,且无需改变底层模型。然而,目前仍不清楚何时进一步的进化会有帮助、成功与失败轨迹如何影响修订,以及额外的测试时计算能否恢复相同的收益。为解决这些问题,我们提出了一个涵盖五个基准和三个模型的受控评估框架。我们的主要研究包含14种支持的模型-基准设置下的42次反馈运行。在每种设置中,我们固定执行器和优化器配置、修订流程、验证规则和轮次预算,仅改变提供给优化器的反馈:成功与失败(Normal)、仅失败或仅成功。进化是稀疏的:388个候选中仅有55个建立了字节级不同的验证最优值。基于验证的选择在14种设置中的11种选择了进化后的技能,其中9种提升了发布测试的性能。所有11次选择都来自包含失败轨迹的反馈条件,尽管Normal和仅失败的相对排名在不同设置中有所不同。在测试、鲁棒性和迁移上的验证与下游评估有时倾向于不同的反馈视图。涵盖八个模型的更广泛的SearchQA分析显示出类似的稀疏、依赖反馈的动态。在GPT-5.5的测试时缩放控制中,Oracle Parallel Sampling与进化后的SearchQA技能的差距在0.43个点以内,但在SpreadsheetBench上仍落后30.96个点;Sequential Refinement则未恢复任何收益。总体而言,持久的技能自进化应被更好地理解为稀疏的、经验证过滤的搜索,其收益依赖于模型和基准,而非通过额外轮次实现的稳定改进。实现代码可在此httpsURL获取。

英文摘要

Self-evolving skill systems promise to improve agents by turning execution feedback into persistent skill updates without changing the underlying model. Yet it remains unclear when further evolution helps, how successful and failed trajectories shape revision, and whether extra test-time computation can recover the same gains. To address these questions, we present a controlled evaluation framework across five benchmarks and three models. Our primary study contains 42 feedback runs across 14 supported model-benchmark settings. Within each setting, we hold the executor and optimizer configuration, revision procedure, validation rule, and round budget fixed, while varying only the feedback shown to the optimizer: successes and failures (Normal), failures only, or successes only. Evolution is sparse: only 55 of 388 candidates establish byte-distinct validation bests. Validation-based selection chooses an evolved skill in 11 of 14 settings, nine of which improve released-test performance. All 11 selections come from feedback conditions that include failed trajectories, although the relative ranking of Normal and Fail-only varies across settings. Validation and downstream evaluations on test, robustness, and transfer sometimes favor different feedback views. A broader SearchQA analysis covering eight models shows similarly sparse, feedback-dependent dynamics. In the GPT-5.5 test-time-scaling controls, oracle Parallel Sampling comes within 0.43 points of the evolved SearchQA skill but remains 30.96 points behind on SpreadsheetBench; Sequential Refinement recovers neither gain. Overall, persistent skill self-evolution is better understood as sparse, validation-filtered search with model- and benchmark-dependent returns, rather than steady improvement from additional rounds. The implementation is available at https://github.com/HKUST-KnowComp/rethinkskill.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑