用于可解释人工智能中鲁棒性与保真度审计的形式化方法论框架:从应用到信任认证
A Formal Methodological Framework for Auditing Robustness and Fidelity in Explainable AI: From Application to Trust Certification
- Doctoral School Modeling–Computer Science, University of Fianarantsoa(菲亚纳兰楚阿大学建模与计算机科学博士学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究构建了审计可解释人工智能(XAI)事后解释器鲁棒性与保真度的形式化框架,以马达加斯加多领域数据集验证发现:高AUC模型或产生无意义解释,审计XAI输出对敏感领域决策至关重要。
AI中文摘要:
SHAP和LIME现已成为解释黑盒预测的标准工具,但当输入被少量噪声扰动时,它们的输出会出现大幅波动——这一问题是我们在之前关于马达加斯加粮食安全的研究(Ralinirina等人,2025)中亲身观察到的。这种可变性引发了一个疑问:这类解释是否完全可信?我们通过构建一个审计协议来解决该问题,该协议可衡量任何事后解释器的两项属性:鲁棒性(输入扰动下解释的稳定性)和保真度(被判定为重要的特征是否实际驱动模型预测)。这两个量被合并为单一的信任分数。我们使用包含83个特征、253条记录和4个营养不良类别的马达加斯加多领域数据集,对3个分类器、2个解释器及其正则化对应模型运行该协议。结果令人警醒:AUC超过0.99的模型可能产生数值退化或完全无意义的解释,且当模型过拟合时,保真度分数会丧失判别能力。这些发现表明,审计XAI输出并非可选,而是必要之举,尤其当这些解释为敏感领域的决策提供依据时。
英文摘要:
SHAP and LIME are now standard tools for interpreting black-box predictions, yet their outputs can vary substantially when the input is perturbed by small amounts of noise--a problem we observed firsthand in our previous work on food security in Madagascar (Ralinirina et al., 2025). This variability raises the question of whether such explanations can be trusted at all. We address it by constructing an auditing protocol that measures two properties of any post-hoc explainer: robustness (how stable the explanation is under input perturbation) and fidelity (whether the features deemed important actually drive the model's prediction). These two quantities are combined into a single Trust Score. We run the protocol on a multi-sectoral dataset from Madagascar (83 features, 253 records, 4 malnutrition classes) using three classifiers and two explainers, plus their regularized counterparts. The results are sobering: models with AUC above 0.99 can produce numerically degenerate or flatly uninformative explanations, and fidelity scores lose discriminative power when the model is overfitted. These findings suggest that auditing XAI outputs is not optional but necessary, particularly when they inform decisions in sensitive domains.