从不确定性到失效归因:分布偏移下用于失效归因的自诊断模型
From Uncertainty to Failure Attribution: Self-Diagnosing Models for Failure Attribution under Distribution Shift
浏览论文内容
中文总结 AI 辅助
针对分布偏移下机器学习模型鲁棒性问题,提出自诊断模型,联合学习预测输出、不确定性及失效归因信号,构建含预定义偏移机制的基准以评估模型失效归因能力。
中文摘要 AI 辅助
分布偏移对机器学习模型的鲁棒性构成重大挑战,但当前解决方案仅旨在检测分布外(OOD)样本并预测不确定性水平。本文提出了分布偏移下失效归因的问题设置,使模型不仅能检测分布外样本,还能找出失效原因。所提出的解决方案称为自诊断模型,能够联合学习预测输出、预测不确定性和失效归因信号。具体而言,我们使用神经网络生成的失效归因向量,该向量通过区分四类不同失效(协方差偏移、语义偏移、噪声损坏和对抗扰动),提供预测不可靠性的结构化表示,即从标量不确定性转向失效识别。为训练模型,我们引入了一致性正则化器,鼓励不确定性与失效归因预测之间的一致性。此外,为评估模型查找失效原因的能力,我们构建了多个带有预定义分布偏移生成机制的分布偏移基准。
英文摘要
Distribution shift poses a significant challenge to the robustness of machine learning models, but the current solutions only aim to detect out-of-distribution (OOD) samples and predict uncertainty levels. We introduce a problem setting for failure attribution under distribution shift, which enables the models not only to detect OOD samples, but also to find out the reason for their failure. The solution we propose is called self-diagnosing models, which are capable of jointly learning predictive output, predictive uncertainty, and a failure attribution signal. In particular, we use the failure attribution vector, produced by a neural network, which provides a structured representation of predictive unreliability by distinguishing four different types of failures: covariance shift, semantic shift, noise corruption, and adversarial perturbation. In other words, we move from scalar uncertainty towards failure identification. For training the model, we introduce a consistency regularizer that encourages consistency between uncertainty and failure attribution predictions. Moreover, to be able to evaluate the model on its ability to find the reasons for failure, we construct several distribution shift benchmarks with predefined mechanisms for generating distribution shifts.