ExplainGuard:一种用于黑盒XAI模型事后解释完整性保证的零信任框架
ExplainGuard: A Zero Trust Framework for Post-Hoc Explanation Integrity Guarantees in Blackbox XAI Models
浏览论文内容
中文总结 AI 辅助
本文提出ExplainGuard零信任框架,通过策略决策点的三项验证支柱,中和XAI解释操纵攻击,将审计过程转为可验证操作,保障黑盒XAI模型事后解释的完整性。
中文摘要 AI 辅助
随着机器学习(ML)模型越来越多地部署在高风险环境中,SHAP和LIME等可解释AI(XAI)方法已成为合规性和建立信任的关键。然而,当前的审计范式依赖于隐含的“信任链”假设,即第三方审计方被视为可信的。近期研究表明,这一假设存在缺陷,对抗性审计方可通过输出洗牌或搭建分布外(OOD)样本等操纵攻击来操控XAI解释,以掩盖模型偏差,同时保持高预测准确率,从而实现“公平洗白”的解释。本文提出一种新型防御框架ExplainGuard,它利用零信任架构(ZTA)设计,可集成到XAI解释供应链中,确保生成解释的完整性。该框架将模糊的“审计方可信”默认假设替换为持续的“先验证再信任”方法。其架构建立了一个策略决策点(PDP),在任何解释向用户发布前,执行三个不同的验证支柱:(1)通过行为指纹实现资产完整性,以检测模型替换;(2)通过公理一致性检查实现语义有效性,以拒绝数学上不可能的解释;(3)利用具有最小计算开销的排名稳定性方法实现特征忠实性验证。最后,我们评估了ExplainGuard如何有效中和最先进的解释操纵攻击,同时将审计过程转变为可验证的操作。
英文摘要
As machine learning (ML) models are increasingly deployed in high-stakes environments, explainable AI (XAI) methods like SHAP and LIME have become essential for regulatory compliance and trust. However, the current auditing paradigm relies on an implicit "chain of trust" where third-party auditors are assumed to be trusted. Recent research demonstrates that this assumption is flawed and adversarial auditors can manipulate XAI explanations through manipulation attacks such as output shuffling or scaffolding out-of-distribution (OOD) to conceal model biases while maintaining high prediction accuracy aiming for fairwashed explanation. In this paper, we introduce a novel defense framework, ExplainGuard, that leverages a Zero-Trust architecture (ZTA) design to be incorporated within the XAI explanation supply chain and ensures the integrity of the generated explanation. This framework would help us to replace the ambiguous default assumption of "auditor is trustworthy," with a continuous "verify-then-trust" approach. Our design architecture establishes a Policy Decision Point (PDP) that enforces three distinct pillars of verification before any explanation is released to the user: (1) asset integrity via behavioral fingerprint to detect model substitution, (2) semantic validity using axiomatic consistency checks to reject mathematically impossible explanations, and (3) feature faithfulness verification utilizing a ranking stability approach with minimal computational overhead. Finally, we evaluate how ExplainGuard can effectively neutralize state- of-the-art explanation manipulation attacks while transforming the auditing process into a verifiable operation.
发表机构
- Tennessee Tech University(田纳西理工大学)
机构由 AI 辅助整理,请以论文原文为准。