arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Task-CoEvolve:通过自适应验证任务选择实现高效的智能体框架优化

Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection

Atsuyuki Miyai, Kiyoharu Aizawa, Toshihiko Yamasaki

arXiv 2608.20169首次发表:更新:

发表机构

The University of Tokyo(东京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Task-CoEvolve通过自适应验证任务选择,使LLM智能体框架与验证任务协同演化,在保持全集搜索最终性能的同时,减少80%评估次数,实现高效的智能体框架优化。

AI 中文摘要

我们提出了一种通过自适应验证任务选择实现高效大语言模型(LLM)智能体框架优化的新方法。框架优化会基于验证性能迭代重写框架代码,无需更新底层模型权重即可实现显著的性能提升。然而,现有方法在每次迭代时都会完整评估固定的验证集,即便随着框架的演化,部分任务的区分度会降低,仍会产生大量评估成本。我们提出了Task-CoEvolve,该方法通过解决两个挑战——选择有信息的任务以及从部分评估中估计全集性能——来使验证任务与框架协同演化。Task-CoEvolve基于以下观察:候选框架存在分歧的任务,相较于那些始终被解决或始终失败的任务,更有助于区分不同的候选框架。它使用基于过往结果的方差加权采样,将评估重点放在智能体能力边界附近的任务上,且采样分布会随框架的演化而调整。随后,它通过考虑采样任务的概率,从采样任务中估计全集分数,即便在评估不同子集的情况下,也能实现跨迭代的一致比较。在在线文本分类和Terminal-Bench 2.1上开展的实验表明,Task-CoEvolve始终优于固定子集基线,且达到了全集搜索的最终性能,同时将优化过程中的评估次数减少了80%。代码将在此httpsURL发布。

英文摘要

We present a novel approach to efficient LLM harness optimization through adaptive validation task selection. Harness optimization iteratively rewrites the harness code based on validation performance, enabling substantial performance gains without updating the underlying model weights. Existing approaches, however, evaluate a fixed validation set in full at every iteration, incurring substantial evaluation costs even on tasks that become less discriminative as the harness evolves. We propose $\textbf{Task-CoEvolve}$, which co-evolves the validation tasks with the harness by addressing two challenges: selecting informative tasks and estimating full-set performance from partial evaluations. Task-CoEvolve builds on the observation that tasks on which candidate harnesses disagree are more informative for distinguishing among them than tasks that are consistently solved or failed. It uses variance-weighted sampling based on past outcomes to focus evaluation on tasks near the capability frontier, with the sampling distribution adapting as the harness evolves. It then estimates full-set scores from the sampled tasks by accounting for their sampling probabilities, enabling consistent comparisons across iterations despite evaluating different subsets. Experiments on online text classification and Terminal-Bench 2.1 show that Task-CoEvolve consistently outperforms subset-based baselines and matches the final performance of full-set search while reducing the number of evaluations during optimization by 80%. Code will be released at https://github.com/Agent4Science-UTokyo/Task-CoEvolve.

Commentsv2: Fix typo in v1, Github: https://github.com/Agent4Science-UTokyo/Task-CoEvolve

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑