arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

根因分析失效之处:检索-重排序分解

Where Root Cause Analysis Fails: A Retrieval-Reranking Decomposition

Hada Melino Muhammad, Luan Pham, Laure Barrière, Sachin Shetty, Leonardo Pulga, Flora D. Salim

arXiv 2609.36686首次发表:更新:

发表机构

University of New South Wales; Baker Hughes(新南威尔士大学; 贝克休斯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出检索-重排序分解,揭示RCA评估盲点,并构建两阶段流水线,在无需因果图或标签数据下,于六个基准上匹配或超越最佳基线top@1准确率。

AI 中文摘要

在复杂监控系统中,从数百个传感器中识别异常的根因对于防止安全事故和代价高昂的停机至关重要。现有研究使用top@k准确率评估根因分析(RCA)方法。我们表明,该指标存在一个根本性盲点:它混淆了两种失败模式,即检索失败(真实原因从未被考虑)和重排序失败(真实原因被考虑但排名过低)。在本工作中,我们引入了一种检索-重排序分解,并审计了四个知名基准以揭示这一盲点。我们的实验表明,在具有复杂故障的基准上,统计基线在79-100%的情况下错误地对真实原因进行排名,而基于图的方法从未明确优于最佳统计基线,无论其因果图是在短故障窗口、在保证包含原因的检索候选池上学习,还是在多天正常运行数据上学习。同时,在故障在其源头显著显现的简单基准上,检索几乎已解决(98-100%)。在分解的指导下,我们构建了一个两阶段流水线,结合多信号检索器与LLM重排序器,作为一个固定配置,在所有六个基准套件上达到或超过最佳基线的top@1准确率(最高提升+12个百分点),无需因果图或标记数据。当所有方法对相同的检索候选进行排名且真实原因保证存在时,添加一个简短的系统描述文档使重排序器在每个基准上领先最佳基线+7到+18个百分点。代码可在该https URL获取。

英文摘要

Identifying the root cause of an anomaly among hundreds of sensors is critical for preventing safety incidents and costly downtime in complex monitored systems. Existing studies evaluate root cause analysis (RCA) methods using top@k accuracy. We show that this metric has a fundamental blind spot: it conflates two failure modes, retrieval failure, where the true cause is never considered, and reranking failure, where it is considered but ranked too low. In this work, we introduce a retrieval-reranking decomposition and audit four well-known benchmarks to expose this blind spot. Our experiments show that, on benchmarks with complex faults, statistical baselines mis-rank the true cause 79-100% of the time, and graph-based methods never clearly beat the best statistical baseline, whether their causal graphs are learned on short fault windows, on retrieved candidate pools guaranteed to contain the cause, or on multi-day normal-operation data. Meanwhile, on simple benchmarks where faults manifest significantly at their origin, retrieval is nearly solved (98-100%). Guided by the decomposition, we build a two-stage pipeline combining a multi-signal retriever with an LLM reranker that, as one fixed configuration, matches or exceeds the best baseline's top@1 accuracy on all six benchmark suites (by up to +12 points), with no causal graph or labeled data required. When all methods rank the same retrieved candidates with the true cause guaranteed present, adding a short system-description document lets the reranker lead the best baseline by +7 to +18 points on every benchmark. Code is available at https://github.com/cruiseresearchgroup/DecompRCA.

CommentsAccepted at NeurIPS 2026 (Evaluations & Datasets Track)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑