arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ORCA-bench:语言模型智能体的值班待命准备程度如何?

ORCA-bench: How Ready Are Language Model Agents for Oncall?

Albert Gong, Kyuseong Choi, Abhineet Agarwal, Jason Schechner, Ryan Huang, Raj Agrawal, Anish Agarwal, Raaz Dwivedi

arXiv 2607.28545首次发表:更新:

发表机构

Cornell Tech; Traversal; Columbia University(康奈尔科技学院; Traversal; 哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究推出ORCA-bench基准,评估编码智能体在生产级值班待命场景下的根本原因分析能力,发现现有前沿智能体表现不佳,凸显其距离安全胜任生产可靠性工作仍有较大差距。

AI 中文摘要

大型语言模型能够编写、修补和搜索代码,但值班待命(oncall)的根本原因分析(RCA)需要不同的能力:从模糊的用户报告出发,对有噪声的指标、日志、追踪和源代码进行推理,且通常要在事件发生数小时后开展。我们推出ORCA-bench,这是一个将通用编码智能体置于生产级保真度值班待命场景中的基准。ORCA-bench将一个实时的、通过OpenTelemetry检测的微服务系统配对,该系统通过Grafana暴露6天的指标、日志和追踪(涉及Prometheus、Jaeger和OpenSearch),并提供完整的源代码访问权限,同时包含1079项RCA任务,这些任务在报告特异性、检测时间和并发故障场景方面存在系统性差异。真实症状由站点可靠性工程(SRE)专家策划并确认,我们的大模型作为评判者的结果已由人工独立重新评分(Cohen's κw=0.90)。在五个前沿智能体中,最佳RCA准确率在中等难度任务(现实输入场景)上为25.3%,在高难度任务上为10.0%——即使使用Claude Fable 5,这一差距仍然存在。最弱的模型在40%的事件报告中会生成不合理的根本原因,而移除源代码访问权限会降低所有指标。关键的是,这些性能是在一个策划好的50GB、6天的测试平台上获得的,其中的任务是在代码和检测工具公开的系统上单独研究的。由于真实生产系统的规模、动态性和特殊性要大几个数量级,我们报告的差距是在前沿编码智能体能够被安全委托负责生产可靠性之前所需工程投入的下限。我们在该httpsURL发布了公开数据集。

英文摘要

Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports, often hours after the incident began. We introduce ORCA-bench, a benchmark that puts general-purpose coding agents in a production-fidelity oncall setting. ORCA-bench pairs 1,079 RCA tasks with six days of metrics, logs, and traces collected from an OpenTelemetry-instrumented microservice system under continuous simulated user load. Agents investigate this recorded history through real observability interfaces---Prometheus, Jaeger, and OpenSearch via Grafana---with full access to application source code. Tasks systematically vary report specificity, time-to-detection, and co-occurring fault scenarios. Ground-truth symptoms are curated and signed off by expert SREs, and our LLM-as-judge is independently re-scored by humans (Cohen's $κ_w = 0.91$). Across five frontier agents, the best RCA Accuracy is 25.3% on Medium-difficulty tasks (the realistic-input setting) and 10.0% on Hard---a gap that remains even with Claude Fable 5. The weakest model hallucinates an implausible root cause in 40% of incident reports, and removing source-code access reduces RCA accuracy and increases the hallucination rate for every evaluated model. These results come from a curated 50 GB / six-day testbed of standalone tasks on a system whose code and instrumentation are public. Since real production systems are orders of magnitude larger, more dynamic, and more idiosyncratic, the gap we report underscores the engineering work still needed before agents can be entrusted with production reliability. We release the public set at https://hub.harborframework.com/datasets/orca-bench/orca-bench.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑