AI 中文总结
本文研究数据漂移下解释保真度是否衰减,发现协变量漂移存在但概念漂移不显著,旧解释不如新解释忠实,需单独监控解释保真度。
AI 中文摘要
模型性能监控是机器学习部署中的标准实践。检测性能被持续跟踪,并且随着特征与目标变量之间的关系退化(这种现象被称为概念漂移),模型衰减是预期的。然而,解释保真度很少以同样的纪律进行监控,即使在金融系统、医疗保健和其他受监管环境等需要解释以用于治理目的的领域中也是如此。本文研究了在数据漂移下,即使特征-目标关系保持稳定,解释是否也会衰减,以及漂移发生前生成的解释是否仍然忠实于替代它们的模型的决策。使用IEEE-CIS交易欺诈检测数据集,我们发现了统计上显著的协变量漂移,但在所实施的条件漂移测试下,没有统计上显著的概念漂移证据,从而提供了一个实证环境,在该环境中,输入分布变化可以与特征-目标关系的可检测变化分开研究。局部解释使用ExIFFI生成,并在三个层面进行评估:路径有效性、结构行为和受控干预下的保真度。结果表明,虽然先前的解释对重新训练的模型保留了实质性的决策相关性,但它们始终不如新生成的解释忠实,并且在所评估的时间窗口内没有系统性扩大的差距的证据。该研究表明,解释保真度需要自己的监控,解释的结构稳定性并不能保证功能保真度,并且解释应被视为与生成它们的模型相关联的工件。
英文摘要
Model performance monitoring is a standard practice in machine learning deployments. Detection performance is tracked continuously, and model decay is expected as the relationship between the feature and target variables degrades, a phenomenon known as concept drift. Explanation fidelity, however, is rarely monitored with the same discipline, even in domains such as financial systems, healthcare, and other regulated environments where explanations are required for governance purposes. This paper investigates whether explanations can decay under data drift, even when the feature-target relationship remains stable, and whether explanations produced before drift occurs remain faithful to the decisions of the model that replaces them. Using the IEEE-CIS Transaction Fraud Detection dataset, we find statistically significant covariate shift but no statistically significant evidence of concept drift under the implemented conditional-drift tests, thereby providing an empirical setting in which input distributional change can be studied separately from detectable changes in the feature-target relationship. Local explanations are generated with ExIFFI and evaluated at three levels: path validity, structural behaviour, and fidelity under controlled intervention. Results show that while prior explanations retain substantial decision relevance to a retrained model, they are consistently less faithful than newly generated explanations, with no evidence of a systematically widening gap across the evaluated windows. The study shows that explanation fidelity requires its own monitoring, that structural stability of explanations does not guarantee functional fidelity, and that explanations should be treated as artifacts tied to the model that produced them.