arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13417cs.AI

超越最终得分:面向长期人工智能研究与开发的智能体系统评估

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

  • Meituan(美团)
  • University of Chinese Academy of Sciences(中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

Yiwei Li, Wanli Yang, Hexiang Tan, Xiangzhou Huang, Zhengyu Chen, Ziran Li, Borun Chen, Shanglin Lei, Huaisheng Zhu, Hao Tian, Fei Sun, Xunliang Cai, Jingang Wang

AI总结:

本文通过新框架评估7个前沿模型在36项长期任务上的表现,发现当前智能体更像工程优化器,方案多基于已有技术,为模型训练等环节改进指明方向。

AI中文摘要:

自主智能体正日益具备通过长期实验改进模型、系统及其他技术产物的能力。然而,为理解该能力的当前状态,评估必须超越最终得分——最终得分既无法揭示进步或损失的环节,也无法表明积累的经验是否能改善后续决策。因此,我们基于一套新框架对7个前沿模型在36项长期任务上开展系统评估,该框架采用基于规则的指标,通过方案构建、执行与反馈控制刻画运行中的行为,并通过受控比较评估任务内及跨任务的经验复用情况。结果显示,当前智能体更像工程优化器而非完全自主的研究者:它们能制定并实施实用方案,但各次运行间性能差异显著;其最优方案主要适配或组合已确立的技术,真正的方法创新仍罕见。详细分析表明,观测到的性能受多重因素影响,包括相似最终结果背后的不同过程瓶颈、可能助力或误导后续决策的经验复用,以及影响性能稳定性的控制框架设计。这些发现为改进模型训练、推理时策略、经验管理及控制框架设计指明了具体方向。

英文摘要:

Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go beyond final scores, which neither reveal where progress is gained or lost nor indicate whether accumulated experience improves later decisions. We therefore present a systematic evaluation of seven frontier models on 36 long-horizon tasks based on a new framework that uses rule-based metrics to characterize within-run behavior through Solution Framing, Execution, and Feedback Control and controlled comparisons to assess experience reuse within and across tasks. The results show that current agents operate more like engineering optimizers than fully autonomous researchers: they can formulate and implement practical solutions, but their performance varies substantially across runs, their strongest solutions mainly adapt or combine established techniques, and genuine methodological novelty remains rare. Detailed analysis reveals that observed performance is shaped by multiple factors, including distinct process bottlenecks behind similar final outcomes, experience reuse that can help or mislead subsequent decisions, and harness designs that affect performance stability. These findings suggest concrete directions for improving model training, inference-time strategies, experience management, and harness design.

↑