arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27069cs.SEcs.LG

图结构在微服务根因分析中是否名副其实?RCAEval 上的受控研究,以及基准真正衡量的是什么

Does Graph Structure Earn Its Place in Microservice Root-Cause Analysis? A Controlled Study on RCAEval, and What the Benchmark Was Really Measuring

  • Polytechnic Faculty, University of Zenica(泽尼察大学理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Imad Buljić

AI总结:

本研究在RCAEval基准上通过受控实验发现图结构对微服务根因分析无可靠增益,揭示基准存在遥测先验偏差,并提出新模型PSC-GRCA及十二项消融研究检查清单。

AI中文摘要:

图神经网络主导了近期微服务根因分析的研究,但最近的结果质疑图结构是否有所贡献。这些结果比较的是整个流水线,因此当平面模型获胜时,无法判断结构是无用的还是冗余的。我们在 RCAEval 上运行了它们所暗示的比较:三个学习分支具有相同的特征、优化器、验证集划分、早停规则和评分头,其中图分支仅由一个术语区分。跨两个 RCAEval 基准、两个拓扑来源和四种机制,我们未发现可靠的图特定效应:在分布内,图模型比传统平面模型领先 0.003 Avg@5(p = 0.844,n = 6 个不相交折叠)。审计流水线揭示了两个基准属性,它们决定了任何结果。RCAEval 仅向每个系统的五个服务注入故障,同时在遥测中暴露 12 到 70 个服务,而主要指标是 Avg@5:一个完全不读取遥测的排序器在 99.7% 的保留事件中将真实故障源排在 top 5 中,Avg@5 得分为 0.488。这个先验,而非均匀随机的 0.137,才是诚实的分布内下限,而它跨系统时降至 0.192。第二个属性是非均匀列模式,它静默地将大多数 RE1 案例的遥测置零。我们复现了已发表的基线 BARO,即 RCAEval 自身的参考实现;在具有干净模式的唯一系统上,它在不同的评分规则下达到了与我们启发式方法相似的聚合准确率,相差在 0.004 以内。审计催生了一个新模型。PSC-GRCA 将候选分数分解为系统先验、遥测证据和中心化图残差,其平均 Avg@5 达到 0.915,而平面基线为 0.864,而其消融研究将大部分增益归因于先验项而非图结构。最后,我们总结了一份包含十二个条目的图与平面消融研究检查清单,这些条目源自记录的六十二个缺陷。

英文摘要:

Graph neural networks dominate recent work on microservice root-cause analysis, yet recent results question whether the graph contributes. Those results compare whole pipelines, so when a flat model wins one cannot tell whether structure is useless or redundant. We run the comparison they imply on RCAEval: three learned arms with identical features, optimiser, validation split, early-stopping rule and scoring head, in which a single term separates the graph arms. Across two RCAEval benchmarks, two topology sources and four regimes we find no reliable graph-specific effect: in-distribution the graph model leads the conventional flat model by 0.003 Avg@5 (p = 0.844, n = 6 disjoint folds). Auditing the pipeline surfaced two benchmark properties that condition any result on it. RCAEval injects faults into only five services per system while exposing 12 to 70 in telemetry, and the headline metric is Avg@5: a ranker reading no telemetry at all places the true culprit in the top five on 99.7 percent of held-out incidents, scoring Avg@5 0.488. That prior, not the uniform-random 0.137, is the honest in-distribution floor, and it collapses to 0.192 across systems. The second property is a non-uniform column schema that silently zeroes telemetry for most RE1 cases. We reproduce a published baseline, BARO, RCAEval's own reference implementation; on the one system with a clean schema it reaches similar aggregate accuracy to our heuristic, within 0.004, under a different scoring rule. The audit motivated a new model. PSC-GRCA separates a candidate score into a system prior, telemetry evidence and a centred graph residual, and reaches mean Avg@5 0.915 against 0.864 for the flat baseline, while its ablations locate most of the gain in the prior term rather than the graph. We close with a twelve-item checklist for graph-versus-flat ablation studies, distilled from sixty-two recorded defects.

补充信息

↑