arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HarnessDev:大型语言模型能否创建并迭代优化自身的智能体测试框架?

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

Yuhao Wu, Jingyuan Zhang, Jiajun Shi, Xinping Lei, Qingshui Gu, Yuxuan Zhang, Zexuan Wang, Chen He, Chen Huang, Maojia Song, Zhiyuan Zeng, Shaowen Wang, Jinkai Liu, Yunfeng Shi, Jiaheng Liu, Shen Yan, Wenhao Huang, Ge Zhang, Wenxuan Zhang

arXiv 2609.01437首次发表:更新:

发表机构

Singapore University of Technology and Design; Georgia Institute of Technology(新加坡科技设计大学; 佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出HarnessDev基准,评估大型语言模型创建和迭代优化自身智能体测试框架的能力,发现其在部分任务上可达到人工基准水平,但性能提升不稳定且跨模型迁移有限。

AI 中文摘要

随着智能体从研究原型发展为部署工具,其能力越来越依赖于模型外部的执行基础设施,通常被称为智能体测试框架(agent harness)。在保持模型权重固定的情况下改变该测试框架,可显著调整任务性能。当前的智能体评估通常报告在选定测试框架下的下游任务性能,而模型自主开发测试框架的能力则相对未被充分探索。我们提出HarnessDev,这一基准将评估单元从任务输出转向可运行的基础设施。HarnessDev涵盖两个阶段:在创建阶段,智能体从最小种子和少量示例出发,构建完整的执行系统;在迭代优化阶段,智能体从自身创建的测试框架出发,利用下游执行反馈进行迭代修改,目标是提升基准性能。随后,我们从能力(在保留基准上的任务成功率)和效率(执行令牌成本)两个维度评估每个构建的测试框架。报告的创建阶段结果涵盖6种创建者大型语言模型、4个领域和5个下游基准,总计2207个独特的下游实例,且开发过程中保留了隐藏的评估任务。我们发现,生成的测试框架在代码、搜索及研究任务上仍显著落后于成熟的人工设计基准,而在写作和机器学习实验任务上与选定基准相当或超越,且执行成本存在较大差异。迭代优化阶段会产生一定性能提升,但这些提升不稳定且仅部分迁移至保留任务。针对固定运行时模型的实验进一步表明,性能提升强烈依赖于执行测试框架的模型,表明跨模型的迁移能力有限。

英文摘要

As agents move from research prototypes to deployed tools, their capability increasingly depends on model-external execution infrastructure, commonly termed the agent harness. Changing this harness while holding model weights fixed can substantially alter task performance. Current agent evaluations typically report downstream performance under a chosen harness, leaving a model's ability to develop the harness itself comparatively underexplored. We introduce HarnessDev, a benchmark that shifts the unit of evaluation from task outputs to runnable infrastructure. HarnessDev covers two stages. In Creation, the agent starts from a minimal seed and a small number of cases, then builds a complete execution system. In Evolution, it starts from its own created harness and iteratively revises it using downstream execution feedback, with the goal of improving benchmark performance. We then evaluate each constructed harness on capability (task success on held-out benchmarks) and efficiency (execution-token cost). The reported Creation results cover six creator LLMs, four domains, and five downstream benchmarks totaling 2,207 unique downstream instances, with hidden evaluation tasks withheld from development. We find that generated harnesses remain substantially behind mature human-engineered references on code and on search and research, while matching or exceeding the selected references on writing and machine-learning experimentation, with large variation in execution cost. Evolution produces some performance gains, but they are unstable and transfer only partially to held-out tasks. Experiments with a fixed runtime model further show that the gains depend strongly on the model executing the harness, indicating limited transfer across models.

CommentsProject page: https://self-developing-agents.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑