arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VERSE:面向智能体框架的验证式自进化优化器

VERSE: Verified Self-Evolving Optimizer for Agent Harnesses

Zekai Wang, Yingqiang Ge, Zekun Wang, Hai Wang, Yuhui Xu, Joshua Frandsen, Shancong Fu, Ashia C. Wilson, Chandan K. Reddy

arXiv 2610.02616首次发表:更新:

发表机构

MIT; Amazon(麻省理工学院; 亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VERSE通过让优化器在提交前测试编辑、重放失败并扰动可疑步骤,实现自进化,在SWE-rebench和分布外任务上显著提升多个框架优化器的性能。

AI 中文摘要

框架进化能够改进LLM智能体的提示、工具和工作流程,而优化器自身的工具和流程往往保持不变。我们研究优化器是否可以通过同时改进其诊断失败、开发编辑和测试效果的方式,更有效地改进另一个智能体。两个观察结果指导了我们的设计。在一项受控研究中,没有基于执行的验证,优化器自进化未能提升性能,但在有验证的情况下取得了该研究的最佳结果。在五个执行器上,自进化优化器构建了用于失败分析、验证、训练审计和工作流控制的自身工具。受这些发现的启发,我们引入了VERSE,一种用于智能体框架的验证式自进化优化器。VERSE允许优化器在提交前测试草稿编辑、重放失败并扰动可疑步骤,同时跨轮次跟踪修复和回归。利用这些反馈,优化器修订执行器框架及其自身的提示、技能、工具、钩子和笔记,而优化器和执行器模型的权重保持不变。在具有不相交训练、验证和测试任务的共享协议下,VERSE在保留的SWE-rebench任务和五种语言的更新分布外任务上改进了所有四个评估的框架优化器。其最佳验证选择的框架分别达到42.3%和37.7%的准确率,而最强基线的准确率分别为39.2%和29.3%。代码可在https://this https URL获取。

英文摘要

Harness evolution improves an LLM agent's prompts, tools, and workflow, while the optimizer's own tools and procedures often remain fixed. We study whether an optimizer can improve another agent more effectively by also improving how it diagnoses failures, develops edits, and tests their effects. Two observations guide our design. In a controlled study, optimizer self-evolution fails to improve performance without execution-based verification, but achieves the best result of that study when verification is available. Across five executors, self-evolving optimizers build their own tools for failure analysis, verification, training audits, and workflow control. Motivated by these findings, we introduce VERSE, a Verified Self-Evolving optimizer for agent harnesses. VERSE lets the optimizer test draft edits, replay failures, and perturb suspected steps before submission, while tracking fixes and regressions across rounds. Using this feedback, the optimizer revises both the executor harness and its own prompts, skills, tools, hooks, and notes, while the weights of the optimizer and executor models stay fixed. Under a shared protocol with disjoint training, validation, and test tasks, VERSE improves all four evaluated harness optimizers on held-out SWE-rebench tasks and newer out-of-distribution tasks in five languages. Its best validation-selected harness reaches 42.3% and 37.7% accuracy, respectively, against 39.2% and 29.3% for the strongest baselines. Code is available at https://github.com/wzekai/VERSE.

Comments45 pages, 13 figures, 15 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑