arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体优化器的效果会叠加吗?基于终端基准测试2.0的持续学习评估

Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0

Wenxiao Wang, Priyatham Kattakinda, Soheil Feizi

arXiv 2607.14004首次发表:更新:

AI 中文总结

研究智能体优化器收益是否叠加的问题,通过基于终端基准测试2.0的两阶段持续学习评估,比较GEPA、Meta Harness和RELAI-VCL三种方法,发现仅RELAI-VCL能使优化收益叠加,因其在优化循环中内置回归控制。

AI 中文摘要

大多数关于智能体优化方法的报告收益都是一次性的:在固定基准上优化智能体,并将由此产生的改进报告为该方法的稳定属性。但这无法测试部署智能体时的实际情况,即随着新故障和新任务出现,优化是递归进行的。核心问题是优化器驱动的收益是否会叠加。我们通过基于终端基准测试2.0中的硬任务构建的两阶段持续学习评估来研究这个问题,在相同优化预算下比较三种智能体优化方法(GEPA、Meta Harness和RELAI的可验证持续学习,即RELAI-VCL)。在传统的静态单阶段设置中,所有三种方法都比基线智能体有所改进。然而,一旦引入新任务,这些方法就出现了明显差异:GEPA优化后的智能体在转移到新任务时表现低于未优化的基线,Meta Harness转移效果良好,但在获得第二次优化预算后无法进一步改进,而RELAI-VCL是唯一一种既能正向转移到未见任务又能在任务纳入优化目标后继续改进的方法,在每个评估阶段都达到最高通过率,总体终身平均通过率最高(GEPA为76.4%,Meta Harness为64.6%,基线为58.7%)。我们的关键观察结果是,只有当优化循环中内置回归控制时,优化收益才会叠加,这为防止无法泛化的捷径解决方案提供了归纳偏差。

英文摘要

Most reported gains from agent-optimization methods are one-shot: an agent is optimized against a fixed benchmark and the resulting improvement is reported as if it were a stable property of the method. This does not test the setting that matters for deployed agents, where optimization is applied recursively as new failures and new tasks appear over time. The central question this raises is whether optimizer-driven gains compound: after an agent has been optimized once, can it be optimized again on newly arrived tasks without eroding the gains the first round produced? We study this question with a two-phase continual-learning evaluation built from hard tasks in Terminal-Bench 2.0, comparing three approaches to agent-harness optimization (GEPA, Meta Harness, and RELAI's Verifiable Continual Learning, RELAI-VCL) under identical optimization budgets. All three methods improve over the baseline agent in the conventional, static, single-phase setting. However, once new tasks are introduced, the methods diverge sharply: GEPA's optimized agent transfers below the unoptimized baseline, Meta Harness transfers well but fails to improve further once given a second optimization budget, and RELAI-VCL is the only method that both transfers positively to unseen tasks and continues improving after those tasks are folded into the optimization objective, reaching the highest pass rate at every evaluated stage and the highest lifelong average pass rate overall (76.4% vs. 66.0% for GEPA, 64.6% for Meta Harness, and 58.7% for the baseline). Our key observation was that optimization gains compounded only when regression control was built into the optimization loop, providing an inductive bias against shortcut solutions that fail to generalize.

CommentsTechnical Report by RELAI (relai.ai)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑