InfraBench:评估跨层级、全生命周期与风险的基础设施智能体
InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk
- University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
- Iowa State University(爱荷华州立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对AI智能体应对真实基础设施复杂度能力不明的问题,提出跨全系统栈、全生命周期且支持细粒度风险评估的InfraBench基准套件,实验表明现有最强智能体仍存在多方面失败模式。
AI中文摘要:
现代计算基础设施因复杂度不断提升,管理难度日益增大。AI智能体的最新进展为基础设施管理任务自动化提供了契机,但这类智能体应对真实基础设施复杂度的能力尚不明确。我们提出InfraBench,这是一个基准套件,用于评估AI智能体在真实基础设施任务中的表现,覆盖全系统栈、全操作生命周期,并支持细粒度风险评估。对15种智能体-模型配置的实验显示,即使是最强的智能体也无法在所有任务中获得满分;平均有效分数范围约为40%至88%(各配置的标准误差为6至12个百分点);每项任务重复执行3次后发现,顶级配置仅能通过部分尝试;按检查项评分则揭示了普遍的失败模式:智能体可能常规性地满足短期目标,却留下非持久化变更、损坏的分布式不变量、不安全的副作用及未清理的状态。INFRABENCH(含实时排行榜、任务与评估工具)可在该http URL公开获取。
英文摘要:
Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity. Recent advances in AI agents create a timely opportunity to automate infrastructure management tasks, but it remains unclear how well such agents can handle real-world infrastructure complexity. We present InfraBench, a benchmark suite for evaluating AI agents on realistic infrastructure tasks across the full system stack and full operational lifecycle with fine-grained risk assessment. Experiments with 15 agent-model configurations show that even the strongest agent cannot secure a full score across all tasks. Mean effective scores range from roughly 40% to 88% (with per-configuration standard errors of 6-12 points), repeating every task three times reveals that top configurations still pass only a fraction of their attempts, and per-check scoring exposes a general failure pattern: agents may routinely satisfy short-term objectives while leaving non-durable changes, broken distributed invariants, unsafe side effects, and uncleaned state behind. INFRABENCH, including its live leaderboard, tasks, and evaluation harness, is publicly available at infraben.ch.