arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37864cs.SEcs.AI

AgentBug-Smith:自动复现智能体系统中的真实世界 Harness 缺陷

AgentBug-Smith: Automatically Reproducing Real-World Harness Bugs in Agentic Systems

Yiming Cheng, Alfin Wijaya Rahardja, Mengshi Zhang, Zihao Chen, Zhenpeng Chen, Yiling Lou

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出 AgentBug-Smith 自动复现智能体系统中的真实 harness 缺陷,并构建 Live-Harness-Bench 基准(含 200 个缺陷),显著提升缺陷修复率,为递归自我改进智能体奠定基础。

中文摘要 AI 辅助

Agent harness 缺陷具有独特性,对最先进的软件智能体而言,修复这些缺陷仍具挑战性。该领域的进展进一步受到现有基准的阻碍,这些基准仅包含少量且固定数量的可执行 harness 缺陷,且构建需要数百个人工小时。本文提出 AgentBug-Smith,一种自动化的 harness 缺陷复现方法,能够从开源智能体系统中持续发现并复现真实世界的 harness 缺陷。在不同的骨干大语言模型(LLM)下,AgentBug-Smith 持续优于为通用软件设计的现有缺陷复现技术,在复现 harness 缺陷的成功率上高出 10.67% 至 27.56%。通过将 AgentBug-Smith 应用于现实中的开源智能体系统,我们构建了 Live-Harness-Bench,一个实时且可扩展的基准,目前包含 200 个可复现的 harness 缺陷。我们进一步通过两个下游应用展示了 Live-Harness-Bench 的实用性。首先,我们使用 Live-Harness-Bench 作为评估基准,系统性地评估最先进的软件智能体,揭示其在修复真实世界 harness 缺陷方面的有限能力。其次,我们使用 Live-Harness-Bench 作为真实世界 harness 缺陷修复的知识库,从中可以提炼出可复用的修复技能,以改进现有软件智能体,将其 harness 缺陷修复率提高 6.32%。总之,AgentBug-Smith 和 Live-Harness-Bench 为持续评估和改进软件智能体的 harness 缺陷修复能力建立了可扩展的基础,将真实世界的智能体故障转化为可执行的评估实例和可复用的 harness 改进知识,从而为递归自我改进的智能体这一终极目标做出贡献。

英文摘要

Agent harness bugs exhibit unique characteristics and remain challenging for state-of-the-art software agents to repair. Progress in this area is further hindered by existing benchmarks, which contain only a small and fixed number of executable harness bugs while requiring hundreds of human hours to construct. This work presents AgentBug-Smith, an automated harness bug reproduction approach that continuously discovers and reproduces real-world harness bugs from open-source agentic systems. Across different backbone LLMs, AgentBug-Smith consistently outperforms existing bug reproduction techniques designed for general software, achieving 10.67% - 27.56% higher success rates of reproducing harness bugs. By applying AgentBug-Smith to open-source agentic systems in the wild, we construct Live-Harness-Bench, a live and extensible benchmark that currently contains 200 reproducible harness bugs. We further demonstrate the utility of Live-Harness-Bench through two downstream applications. First, we use Live-Harness-Bench as the evaluation benchmark to systematically evaluate state-of-the-art software agents, revealing their limited capabilities in repairing real-world harness bugs. Second, we use Live-Harness-Bench as a knowledge base of real-world harness bug fixes, from which reusable repair skills can be distilled to improve existing software agents, increasing their harness-bug repair rates by 6.32%. Together, AgentBug-Smith and Live-Harness-Bench establish a scalable foundation for continuously evaluating and improving software agents on harness bug repair, turning real-world agent failures into executable evaluation instances and reusable knowledge for harness improvement, thus contributing to the ultimate goal of recursively self-improving agents.

发表机构

  • The University of Chicago(芝加哥大学)
  • Fudan University(复旦大学)
  • TensorBlock, Inc.(TensorBlock公司)
  • Tsinghua University(清华大学)
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑