arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25661cs.SEcs.AI

从通用智能体到RCA专家:一种用于根因分析的自演化框架

From General Agents to RCA Experts: A Self-Evolving Harness for Root Cause Analysis

  • The Chinese University of Hong Kong(香港中文大学)
  • ByteDance(字节跳动)

机构由 AI 辅助整理,请以论文原文为准。

Haiyu Huang, Jiewei Lyu, Zhihan Jiang, Jinyang Liu, Xiao He, Tieying Zhang, Wu Xiang, Michael R. Lyu

AI总结:

该研究提出自演化RCA框架OpsHarness,复用通用智能体能力,通过双门验证演化,在基准与工业场景中大幅提升RCA准确率。

AI中文摘要:

基于大语言模型(LLM)的自动化根因分析(RCA)已受到越来越多的关注。如今,站点可靠性工程师(SRE)通常以两种方式之一借助LLM实现RCA自动化:要么直接使用通用智能体(例如Codex或Claude Code)进行诊断,要么从零构建专用RCA智能体。随着主流通用智能体能力不断提升且迭代迅速,我们的定量研究发现,前者如今通常优于后者。然而,其准确率仍未达到生产需求,这一差距主要源于智能体通用能力之外的外部适配层,即框架(harness)。因此,我们认为基于LLM的RCA应聚焦于这一外部框架,复用现代通用智能体的强大通用能力,而非从零重建智能体。此类框架的一项关键能力是自演化,即从过往诊断中积累特定系统的经验,使其使用次数越多就越出色。我们推出OpsHarness,一种自演化RCA框架,可将诊断经验转化为可复用的专业知识。其数据平面将分层的运维知识与创意卡片工具库相结合,而控制平面则协调设置、诊断、演化与验证。演化过程中,OpsHarness会对比成功与失败的轨迹,将其证据转化为原子提案,且仅通过双门验证流程接受更新,该流程旨在防止过拟合与退化。在两个公开基准测试及一次工业部署中,OpsHarness实现了59.0%的top-1准确率,较纯通用智能体提升63.4%,较基线RCA智能体提升4.02倍。

英文摘要:

Automated root cause analysis (RCA) with large language models (LLMs) has drawn growing attention. Today, SREs typically automate RCA with LLMs in one of two ways: directly using a general-purpose agent (e.g., Codex or Claude Code) for diagnosis, or building a specialized RCA agent from scratch. As mainstream general agents grow more capable and iterate quickly, our quantitative study finds that the former now often surpasses the latter. Its accuracy, however, still falls short of production needs, and this gap stems mainly from the external adaptation layer outside the agent's general capabilities, namely the harness. We therefore argue that LLM-based RCA should focus on this external harness, reusing the strong general capabilities of a modern agent rather than rebuilding an agent from scratch. A key capability of such a harness is to self-evolve, accumulating system-specific experience from past diagnoses so that it gets better the more it is used. We introduce OpsHarness, a self-evolving RCA harness that turns diagnosis experience into reusable expertise. Its data plane combines layered operational knowledge with an idea-card tool library, while its control plane coordinates setup, diagnosis, evolution, and verification. During evolution, OpsHarness contrasts successful and failed trajectories, converts their evidence into atomic proposals, and admits updates only through a dual-gate verification process designed to prevent overfitting and regression. Across two public benchmarks and an industrial deployment, OpsHarness achieves 59.0\% top-1 accuracy, improving over a bare general agent by 63.4\% and over baseline RCA agents by 4.02$\times$.

↑