过于自信难安全:面向可靠日志异常检测的模型校准
Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection
- Beijing Jiaotong University(北京交通大学)
- University of Florida(佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对基于语言模型的日志异常检测器置信度校准差的问题,提出轻量级事后校准框架LoRD,经实验验证其可提升置信度可靠性并减少过度自信异常相关错误且不牺牲检测性能。
AI中文摘要:
在线日志异常检测对维持大规模计算系统的可靠性至关重要。尽管近期基于语言模型的日志异常检测器取得了出色的检测性能,但其置信度估计的校准效果仍较差。我们发现这些检测器频繁为错误预测分配过高的置信度,尤其是在类别严重不平衡情况下的异常日志。此外,即使传统校准指标显示校准效果良好,错误预测的置信度仍持续处于高位,这为运营监控系统造成了关键的可靠性缺口。为解决该问题,我们提出Log Reconstruction and Distance(LoRD),一种用于可靠日志异常检测的轻量级事后校准框架。LoRD从正确分类的验证样本的潜在表示中学习特定预测路径的可靠性模型,并通过路径级重建距离估计预测可靠性。基于估计的可靠性,LoRD选择性地重新校准高风险预测,以抑制过度自信的错误,同时保留可靠的预测。在四个大规模日志基准数据集和多个基于语言模型的检测器上进行的大量实验表明,LoRD可持续提升置信度可靠性,并在不牺牲异常检测性能的前提下大幅减少与异常相关的过度自信错误。
英文摘要:
Online log anomaly detection is critical for maintaining the reliability of large-scale computing systems. Although recent language model-based log anomaly detectors achieve strong detection performance, their confidence estimates remain poorly calibrated. We show that these detectors frequently assign excessive confidence to incorrect predictions, particularly for anomalous logs under severe class imbalance. Moreover, confidence on erroneous predictions remains persistently high even when conventional calibration metrics indicate good calibration, creating a critical reliability gap for operational monitoring systems. To address this issue, we propose Log Reconstruction and Distance (LoRD), a lightweight post-hoc calibration framework for reliable log anomaly detection. LoRD learns prediction-route-specific reliability models from latent representations of correctly classified validation samples and estimates prediction reliability through route-wise reconstruction distances. Based on the estimated reliability, LoRD selectively recalibrates high-risk predictions to suppress overconfident errors while preserving reliable predictions. Extensive experiments on four large-scale log benchmark datasets and multiple language model-based detectors demonstrate that LoRD consistently improves confidence reliability and substantially reduces overconfident anomaly-related errors without sacrificing anomaly detection performance.