发表机构
College of Computer Science and Technology, Zhejiang University; School of Software Technology, Zhejiang University; China Telecom Cloud Computing Corporation(浙江大学计算机科学与技术学院; 浙江大学软件学院; 中国电信云计算公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EviRCA将确定性证据提取与LLM推理解耦,通过生成多模态证据卡片和只读工具,在OpenRCA基准上以40.6%-43.9%正确率大幅超越基线,并显著降低计算成本。
AI 中文摘要
根因分析(RCA)是维护现代微服务系统的一项关键但劳动密集型的任务,因此成为大型语言模型(LLM)的一个有吸引力的应用目标。最近的智能体方法允许LLM通过生成和执行代码来迭代探索原始遥测数据,要求单一模型同时在大规模异构遥测数据上检索证据、定位故障并推断根因,这导致了高计算成本、不稳定的行为和有限的诊断准确性。然而,原始遥测数据由数值指标、结构化追踪和机器生成的日志组成,并不直接适合LLM处理。我们提出了EviRCA,一个基于LLM的RCA框架,将确定性证据提取与LLM推理解耦。一个与系统无关的提取阶段将原始指标、追踪和日志转换为一组紧凑的忠实多模态证据卡片,而LLM仅通过一小组预定义的只读工具在这些结构化观察上进行推理,无需访问原始遥测数据或执行代码。我们在OpenRCA上评估了EviRCA,OpenRCA是一个基于三个企业系统的真实异构遥测数据构建的基准。EviRCA在两个不同的LLM上实现了40.6%-43.9%的正确率,大幅优于先前OpenRCA基线(最高达到15.2%),同时将令牌消耗减少了15-26倍,执行时间减少了3-20倍。此外,EviRCA解决了需要同时推理时间、组件和根因的困难案例,而先前的方法在此设置下报告了接近零的性能。我们的过程级失败分析进一步表明,瓶颈在于判断提取阶段已经呈现的证据,而非搜索证据,这表明基于LLM的RCA的有效性在很大程度上取决于证据提取的质量。
英文摘要
Root-cause analysis (RCA) is a critical yet labor-intensive task for maintaining modern microservice systems, making it an attractive target for large language models (LLMs). Recent agentic approaches allow an LLM to iteratively explore raw telemetry by generating and executing code, asking a single model to simultaneously retrieve evidence, localize faults, and infer root causes over large volumes of heterogeneous telemetry, which leads to high computational cost, unstable behavior, and limited diagnostic accuracy. However, raw telemetry consists of numeric metrics, structured traces, and machine-generated logs that are not directly suitable for LLM processing. We present EviRCA, a framework for LLM-based RCA that decouples deterministic evidence extraction from LLM reasoning. A system-agnostic extraction stage converts raw metrics, traces, and logs into a compact set of faithful multimodal evidence cards, while the LLM reasons only over these structured observations through a small set of predefined read-only tools, without accessing raw telemetry or executing code. We evaluate EviRCA on OpenRCA, a benchmark built from real, heterogeneous telemetry across three enterprise systems. EviRCA achieves a correct rate of 40.6%-43.9% across two different LLMs, substantially outperforming prior OpenRCA baselines that achieve up to 15.2%, while reducing token consumption by 15-26x and execution time by 3-20x. Moreover, EviRCA solves hard cases requiring simultaneous reasoning over time, components, and root causes, a setting where previous approaches reported near-zero performance. Our process-level failure analysis further shows that the bottleneck lies in judging the evidence that the extraction stage has already surfaced, rather than searching for it, suggesting that the effectiveness of LLM-based RCA depends heavily on the quality of evidence extraction.
Comments12 pages, 4 figures, 6 tables