arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21310cs.SE

超越故障定位:面向微服务根本原因分析的大语言模型智能体的轨迹级研究

Beyond Fault Localization: A Trajectory-Level Study of LLM Agents for Microservice Root Cause Analysis

Qisheng Lu, Aoyang Fang, Junjielong Xu, Jin'ao Shang, Songhan Zhang, Yifan Yang, Xiaochuan Yan, Pinjia He

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对微服务根本原因分析,提出轨迹级评估框架,发现智能体答案正确性与诊断质量脱节,构建DiagGuard架构提升了RCA的Acc@1指标,为自动化RCA改进提供指导。

中文摘要 AI 辅助

现有针对微服务自动化根本原因分析(RCA)的评估主要通过端点正确性来衡量诊断性能:即方法是否能定位到责任服务。该标准虽便于比较,但无法揭示诊断的证据基础,也无法揭示从故障源到观测症状的故障传播路径,而这两点正是值班站点可靠性工程师判断是否需要采取行动的依据。因此,我们将RCA视为一个可观测的诊断过程。我们的轨迹级框架针对人工整理的服务级故障传播路径评估智能体的执行情况。将该框架应用于公开的微服务RCA基准后,它分析了3500条诊断轨迹,刻画了智能体的调查方向以及如何使用检索到的遥测数据。我们发现答案正确性与诊断质量之间存在脱节:智能体可能能定位故障源,但无法重建其传播过程。成功的调查会保持在故障影响范围内,基于检索到的证据采取行动,并随着搜索深入拓宽查询范围。失败的情况包括遗漏决定性证据、错误解释检索到的证据,或用无依据的推理替代缺失的证据。我们将这种分类方法化为DiagGuard,这是一种两阶段纵深防御架构,其中基于 grounding (基于观测)先调查可用观测数据再进行定位,验证阶段则对照观测数据审核诊断结果。在采用不同模型、基准和服务拓扑的独立环境中,DiagGuard将Acc@1从43.5%提升至52.5%。这些结果表明,轨迹级评估能揭示最终答案指标所隐藏的局限性,并为改进自动化RCA提供可行指导。

英文摘要

Existing evaluations of automated root cause analysis (RCA) for microservices assess diagnostic performance mainly by endpoint correctness: whether a method localizes the responsible service. This criterion enables comparison but does not reveal the evidentiary basis of a diagnosis or the fault-propagation route connecting the source to observed symptoms, both of which an on-call site reliability engineer needs to judge whether action is warranted. We therefore treat RCA as an observable diagnostic process. Our trajectory-level framework evaluates agent executions against manually curated service-level fault-propagation paths. Applied to a public microservice RCA benchmark, it analyzes 3,500 diagnostic trajectories, characterizing where agents investigate and how they use retrieved telemetry. We find a disconnect between answer correctness and diagnostic quality: an agent may localize the fault source yet fail to reconstruct its propagation. Successful investigations stay on the fault-impact surface, act on retrieved evidence, and broaden their query repertoire as the search deepens. Failures arise when decisive evidence is omitted, retrieved evidence is misinterpreted, or unsupported inference substitutes for missing evidence. We operationalize this taxonomy as DiagGuard, a two-stage defense-in-depth architecture in which grounding surveys available observations before localization and verification audits the diagnosis against them. In an independent setting with a different model, benchmark, and service topology, DiagGuard raises Acc@1 from 43.5% to 52.5%. These results show that trajectory-level evaluation exposes limitations hidden by final-answer metrics and provides actionable guidance for improving automated RCA.

补充信息

↑