发表机构
University of Copenhagen(哥本哈根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对机器学习模型的可解释性,提出Ra-NEM方法,通过优化插入和删除曲线下面积损失函数,高效生成高忠实性归因,且不牺牲模型性能。
AI 中文摘要
机器学习在社会相关任务中的采用需要有效的可解释人工智能(XAI)方法来更好地理解机器学习模型的行为。归因方法是一种流行的XAI方法,其中输入-输出关系通过热力图来表征,这些热力图反映了特定预测中输入特征的相对重要性。此类图的质量通常通过基于插入和删除曲线下面积来测量忠实性,该面积测量随着特征的添加和移除而模型输出的变化。在本研究中,我们从这个忠实性概念中推导出一个目标函数,并找到一种近似其梯度的方法。我们建立了插入曲线与top-$k$特征选择之间的联系,这导致了一个衡量归因质量的损失函数。对损失进行随机化使我们能够高效地近似其梯度。为了展示一般方法的有效性,我们将该损失函数与神经解释掩码框架相结合。由此产生的方法,称为Ra-NEM,可以与任何可微模型一起使用,而不影响模型的性能。实验表明,Ra-NEM能够稳健且高效地提供准确的归因。与其他算法相比,归因不仅具有更高的忠实性,而且在其他XAI指标上也表现良好。Ra-NEM的高推理速度使得该方法适用于在线应用。代码可在网上获取:此https URL
英文摘要
The adoption of machine learning for socially relevant tasks requires effective explainable artificial intelligence (XAI) methods to better understand the behavior of machine learning models. Attribution methods are a popular XAI approach in which input-output relationships are characterized by heat maps that reflect the relative importance of input features for a particular prediction. The quality of such maps is often assessed by measuring faithfulness based on the area under insertion and deletion curves, which measures changes in the model output as features are added and removed. In this study, we derive an objective function from this notion of faithfulness and a way to approximate its gradient. We establish the connection between insertion curves and top-$k$ feature selection, which leads to a loss function measuring the quality of attributions. Randomization of the loss allows us to efficiently approximate its gradient. To show the effectiveness of the general approach, we combine the loss function with the neural explanation mask framework. The resulting method, termed Ra-NEM, can be used with any differentiable model without affecting the model's performance. Experiments demonstrate that Ra-NEM provides accurate attributions robustly and efficiently. Compared to other algorithms, the attributions have not only higher faithfulness but also perform well in terms of other XAI metrics. The high inference speed of Ra-NEM makes the method suitable for online applications. The code is available online: https://github.com/baerminator/Ra_Nem