通过互补多视图可解释性理解Linux中的在线故障预测
Understanding Online Failure Prediction in Linux Through Complementary Multi-View Explainability
- University of Coimbra(科英布拉大学)
- CISUC(科英布拉大学信息系统与计算机中心)
- LASI(智能系统实验室)
- Department of Informatics Engineering(信息工程学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文构建并评估Linux的可解释在线故障预测流水线,其跨工作负载检测准确率达91%-94%、误报率低于1%,但诊断对工作负载转移更敏感,留一模式评估下未见故障模式诊断准确率为0%,凸显互补可解释性机制的价值。
AI中文摘要:
准确的在线故障预测(Online Failure Prediction, OFP)已被证明在操作系统(Operating Systems, OSs)环境中是可行的,但仅预测本身不足以实现实际应用。若缺乏诊断洞察,运维人员将难以信任预警或决定如何响应。此外,即使预测准确率很高,通常也不清楚模型是在捕捉有意义的故障过程,还是仅利用了工作负载特有的噪声和遥测中的偶然相关性。本文报告了构建和评估Linux操作系统可解释OFP流水线的实践经验。我们将用于检测的基于共识的特征选择与时序起始分析、子系统级因果分析及互补诊断机制相结合,以支持故障解释。在冻结训练工件的严格跨工作负载条件下评估,该流水线在未重新训练的情况下对未见工作负载实现了91%-94%的检测率,同时将误报率保持在1%以下。然而,故障模式诊断对工作负载转移的敏感度显著更高,且若干诊断机制对特定故障类型的有效性有限。我们的经验凸显了三个主要教训:i)在工作负载变化下,检测比诊断更具鲁棒性;ii)预警能力在很大程度上取决于故障模式,本研究中从38秒到215秒不等;iii)仅从相关训练模式无法可靠诊断未见故障模式,在留一模式(Leave-One-Mode-Out, LOMO)评估下准确率为0%。综上,这些结果表明互补可解释性机制对解释准确故障预测的价值,揭示了预测信号何时反映可迁移的故障结构,以及诊断泛化何时在工作负载变化下失效。
英文摘要:
Accurate Online Failure Prediction (OFP) has been shown to be feasible in Operating Systems (OSs) settings, but prediction alone is not sufficient for practical adoption. Without diagnostic insight, operators have limited basis to trust alerts or decide how to respond. Moreover, even when predictive accuracy is high, it is often unclear whether models are capturing meaningful failure processes or merely exploiting workload-specific noise and incidental correlations in telemetry. This paper reports a practical experience building and evaluating an explainable OFP pipeline for Linux OSs. We combine consensus-based feature selection for detection with temporal onset analysis, subsystemlevel causal analysis, and complementary diagnostic mechanisms to support failure interpretation. Evaluated under strict crossworkload conditions with frozen training artifacts, it achieved 91-94% detection on unseen workloads without retraining, while maintaining false alarm rates below 1%. However, failure mode diagnosis proved substantially more sensitive to workload shift, and several diagnostics mechanisms showed limited effectiveness for specific failure types. Our experience highlights three main lessons: i) detection generalizes more robustly than diagnosis across workload changes; ii) early-warning capability depends strongly on the failure mode, ranging from 38 to 215 seconds in our study; and iii) unseen failure modes are not reliably diagnosable from related training modes alone, providing 0% accuracy under Leave-One-Mode-Out (LOMO) evaluation. Taken together, these results show the value of complementary explainability mechanisms for interpreting accurate failure predictions, revealing when predictive signals reflect transferable failure structure and when diagnostic generalization breaks down under workload variation.