arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Evo-Bench:语言模型能否改进智能体框架?

Evo-Bench: Can Language Models Improve Agent Harness?

Lisheng Huang, Chen Yang, Hao Zhou, Huatong Song, Zongchao Chen, Ran Le, Yang Song, Wayne Xin Zhao, Tao Zhang

arXiv 2608.09096首次发表:更新:

发表机构

Gaoling School of Artificial Intelligence, Renmin University of China; BOSS Zhipin(中国人民大学高瓴人工智能学院; BOSS直聘)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出首个智能体框架进化基准Evo-Bench,评估语言模型的自主框架优化能力,发现其在通用、搜索任务中优于人工框架,在办公任务中表现不佳,且合成框架可迁移提升各类模型。

AI 中文摘要

大型语言模型(LLMs)推动了自主智能体的快速发展,但标准评估仍局限于静态任务解决。一个新兴前沿是框架进化,即智能体自主优化自身操作框架的能力。然而,对这一能力进行系统基准测试仍具挑战,因为现有评估无法将框架改进与基础模型实力分离、无法防止任务特定过拟合,也无法捕捉长周期迭代研究。为应对这些挑战,我们推出Evo-Bench,这是首个旨在评估模型在搜索、办公和通用智能体领域内在框架进化能力的基准。为严格隔离该能力,Evo-Bench采用了一种新颖的框架引导构建框架:它利用辅助任务进化识别对框架改进真正敏感的任务,随后通过感知敏感性的分层拆分确保跨套件的稳健泛化。对9个前沿和开放权重模型的广泛评估显示,顶级模型实现了高达16.6分的巨大绝对提升,接近最先进的人工设计基准。关键的是,自主进化在通用任务中优于人工框架,且在搜索任务中表现出色,但在需要高度特定处理工作流的办公任务中表现挣扎。此外,我们的分析揭示了早期饱和等关键时间异常,同时表明合成的框架作为高度可迁移的推理结构,能持续提升各类策略模型。

英文摘要

Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously optimize its own operating harness. However, systematically benchmarking this capability remains challenging, as existing evaluations fail to isolate harness improvements from base model strength, prevent task-specific overfitting, or capture long-horizon iterative research. To address these challenges, we introduce Evo-Bench, the first benchmark designed to evaluate models' intrinsic harness-evolving capabilities across Search, Office, and General agent domains. To rigorously isolate this capability, Evo-Bench employs a novel harness-guided construction framework: it leverages auxiliary-task evolution to identify tasks genuinely sensitive to framework improvements, followed by sensitivity-aware stratified splitting to ensure robust cross-suite generalization. Extensive evaluations across nine frontier and open-weight models reveal that top models achieve massive absolute gains reaching 16.6 points, closely approaching state-of-the-art human-engineered baselines. Crucially, while autonomous evolution outpeforms artificial harness in General tasks and excels in Search tasks, it struggles in Office tasks that demand highly specific processing workflows. Furthermore, our analysis exposes critical temporal anomalies like early saturation, while demonstrating that the synthesized harnesses act as highly transferable reasoning structures, consistently boosting diverse policy models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑