SRE-Marathon:面向自主站点可靠性代理的持续、变更驱动基准
SRE-Marathon: A Continuous, Change-Driven Benchmark for Autonomous Site Reliability Agents
浏览论文内容
中文总结 AI 辅助
针对现有SRE代理基准的情节式局限,提出持续、变更驱动的SRE-Marathon基准,模拟生产运维,实验表明最佳方法得分仅41.3,代理能关联定位故障但难以完成修复。
中文摘要 AI 辅助
站点可靠性工程(SRE)代理的基准测试通常是情节式的:注入一个故障,代理接收一个事件任务,然后对其响应进行评分。生产运维并非如此。事件通过嘈杂的警报浮出水面,在时间上重叠,并且通常源于代码或配置变更。我们提出了SRE-Marathon,一个面向长时域、持续SRE运维的基准。代理以固定节奏被调用,接收累积的警报历史和持久化工作空间,同时操作一个活跃的双区域Kubernetes部署,故障编排器根据种子化的、生产校准的时间表注入重叠故障。精选的代码和配置变更通过相同的构建管道部署,可用于修复。每次运行被记录到一个密封的捆绑包中,并离线评分:Marathon-Score对每个注入的故障,根据相关性、定位和修复的有序进展进行积分,所有指标均从记录的系统证据中确定性计算。在三个应用和每次运行约六十个故障的情况下,10种方法中最好的仅达到41.3分(满分100)。代理通常能关联和定位故障,但几乎从未在故障活跃期间完成修复。
英文摘要
Benchmarks for site reliability engineering (SRE) agents are typically episodic: one fault is injected, the agent receives an incident task, and its response is scored. Production operation is not. Incidents surface through noisy alerts, overlap in time, and often originate from code or configuration changes. We present SRE-Marathon, a benchmark for long-horizon, continuous SRE operation. An agent is invoked at a fixed cadence with cumulative alert history and a persistent workspace while operating a live two-zone Kubernetes deployment as a fault orchestrator injects overlapping faults according to a seeded, production-calibrated schedule. Curated code and configuration changes deployed through the same build pipeline available for repair. Each run is recorded into a sealed bundle and scored offline: Marathon-Score credits each injected fault for ordered progress through correlation, localization, and repair, with all metrics computed deterministically from recorded system evidence. Across three applications and about sixty faults per run, the best of 10 methods reaches only 41.3 out of 100. Agents often correlate and localize faults, but almost never complete repairs while the faults remain active.