发表机构
Abundant AI; Adrenaline AI; Carnegie Mellon University; Comenius University in Bratislava; Massachusetts General Hospital(Abundant AI; Adrenaline AI; 卡内基梅隆大学; 布拉迪斯拉发夸美纽斯大学; 马萨诸塞总医院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Incident-Arena 是一个基于真实开源软件的人工构建基准,通过 20 个任务和功能验证器评估 AI 智能体在生产事件响应中的可靠性,发现前沿模型得分低于 64.3%,暴露了长程推理和修复安全性的不足。
AI 中文摘要
AI 编码智能体在工业界和学术界的工程工作流中无处不在。然而,尽管它们在应用编码中被广泛使用,但对其在生产事件响应中执行能力的关注相对较少。这一新兴领域被称为智能体站点可靠性工程(SRE),其现有基准受限于(1)不切实际的环境,通常是玩具仓库;(2)非标准的框架实现;(3)简单的静态验证器。我们引入了 Incident-Arena,这是一个人工构建的基准,包含 20 个精心挑选的任务,这些任务基于真实世界部署的开源软件。每个任务将一个生产应用部署到临时 Kubernetes 集群中,从配置层到底层镜像注入故障,并根据任务要求提供持续的负载配置文件。我们还提出了一种新颖的验证方法,超越静态检查,采用功能验证器,保持系统级指标稳定,同时确保修复安全进行。智能体试验平均运行 281 万 token 和 41 轮交互,超越了现有基准,展示了智能体的长程推理能力。在 20 个任务和 3 个应用基座上,前沿模型得分低于 64.3%,失败从诊断/定位错误延伸到不完整修复和不安全回归。
英文摘要
AI coding agents are ubiquitous in engineering workflows amongst industry and academia. Yet, despite their use in app coding, relatively less attention has been paid to their ability to execute on production incident response. This emerging field, termed agentic site-reliability-engineering (SRE) contains benchmarks limited by (1) unrealistic environments, typically toy repositories (2) non-standard framework implementations and (3) simple static verifiers. We introduce Incident-Arena, a human-built benchmark of 20 carefully selected tasks grounded in real-world deployed open source software. Each task deploys a production application to an ephemeral Kubernetes cluster, injecting a fault from the config layer through underlying images, and a sustained load profile given the task requirements. We also present a novel verification method, going beyond static checks to functional verifiers, holding systems level metrics stable, while ensuring repairs are done safely. Agent trials run an average of 2.81M tokens and 41 turns, going beyond existing benchmarks, demonstrating agentic long horizon reasoning. Across 20 tasks and 3 application substrates, frontier models score below 64.3%, with failures extending from diagnosis/localization errors, through incomplete repairs and unsafe regressions.