Bifrost:利用基于日志的故障诊断的可错性表示赋能预训练语言模型
Bifrost: Empowering Pretrained Language Model with Fallibility Representation for Log-Based Fault Diagnosis
浏览论文内容
中文总结 AI 辅助
研究基于日志的故障诊断问题,提出Bifrost方法,借鉴站点可靠性工程师经验,基于自监督对比学习设计策略学习日志可错性表示,在多个系统中其生成的日志表示在故障诊断相关指标上优于现有PLMs。
中文摘要 AI 辅助
基于日志的故障诊断对于运行时调试和维护至关重要。现有故障诊断方法使用在自然语言上预训练的语言模型(PLMs)进行日志表示。然而,系统故障反映在系统日志的多层次结构中。在自然语言上预训练的PLMs难以全面捕获多层次故障信息,无法满足故障诊断的要求。我们将此信息称为可错性表示。为解决此问题,我们提出了一种新颖的日志表示学习方法Bifrost。它从站点可靠性工程师的日志分析经验中汲取灵感,并基于自监督对比学习精心设计策略来学习日志的可错性表示。在三个公共系统和一个工业机器学习即服务系统中,Bifrost生成的日志表示在异常检测的F1中平均比现有PLMs高出9.83%,在根本原因定位的HR@k中高出18.28%,在故障识别的Macro-F1中高出20.88%。
英文摘要
Log-based fault diagnosis is crucial for runtime debugging and maintenance. Existing fault diagnosis methods use language models pre-trained on natural language (PLMs) for log representation. However, system faults are reflected in the multi-level structure of system logs. PLMs pre-trained on natural language struggle to comprehensively capture multi-level fault information, failing to meet the requirements of fault diagnosis. We refer to this information as fallibility representations. To address this problem, we propose a novel log representation learning method, Bifrost. It draws inspiration from the log analysis experience of Site Reliability Engineers and meticulously designs strategies based on self-supervised contrastive learning to learn the fallibility representations of logs. Across three public systems and one industrial ML-as-a-Service system, the log representations produced by Bifrost outperform existing PLMs by average margins of 9.83% in F1 for anomaly detection, 18.28% in HR@k for root cause localization, and 20.88% in Macro-F1 for fault identification.