arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

摧毁证据:一种用于欺骗白盒可解释性AI审计器的双惩罚规避框架

Crushing the Evidence: A Dual-Penalty Evasion Framework for Fooling White-Box Explainable AI Auditors

Niraj Kumar, Harsh Kasyap

arXiv 2608.00566首次发表:更新:

AI 中文总结

该研究提出一种双惩罚白盒规避攻击框架,将规避逻辑嵌入模型参数,可压制目标特征归因至近零、维持高攻击成功率,且能绕过条件异常检测,欺骗可解释性AI审计器。

AI 中文摘要

事后模型解释器(如LIME、SHAP和Integrated Gradients)被广泛部署于金融、医疗、社会福利等高风险敏感领域,以审计模型,确保模型的透明度与可接受性。然而,仅有少量研究探讨了可解释性流程中潜在的攻击:攻击者可通过对抗性解释攻击掩盖算法偏见或后门,这类攻击依赖分布外(OOD)检测器,该检测器在被解释器查询时会切换预测。因此,防御措施已被开发,通过识别此类攻击的异常扰动痕迹,成功中和这些黑盒攻击。本文中,我们引入一种更强大的白盒、梯度正则化规避攻击框架,以揭示一个关键漏洞。通过采用连续嵌入双惩罚框架,我们在分布内数据训练期间直接惩罚触发特征梯度。由于我们的方法将规避逻辑原生嵌入模型参数,不依赖OOD包装器,因此生成平滑的分布内预测,且无异常痕迹。在四个基准表格数据集(COMPAS、German Credit、IEEE-CIS和Communities & Crime)上的实证评估证实,我们的方法系统性地将目标特征归因压制至接近零(<0.02),维持>90%的攻击成功率,并从根本上绕过条件异常检测。

英文摘要

Post-hoc model explainers such as LIME, SHAP, and Integrated Gradients are widely deployed to audit models in high-stakes sensitive domains, including finance, healthcare, and social welfare. This ensures the model's transparency and acceptability. However, a few studies have examined potential attacks in the explainability pipeline. Adversaries can attempt to conceal algorithmic biases or backdoors using adversarial explanation attacks. These attacks have relied on scaffolding out-of-distribution (OOD) detectors that toggle predictions when queried by an explainer. Consequently, defenses have been developed to successfully neutralize these black-box attacks by identifying their anomalous perturbation footprints. In this paper, we demonstrate a critical vulnerability by introducing a more potent white-box, gradient-regularized evasion attack framework. By employing a continuous-embedding dual-penalty framework, we directly penalize trigger feature gradients during training on in-distribution data. Since our approach embeds the evasion logic natively into the model parameters, without relying on OOD scaffolding wrappers, it generates smooth, in-distribution predictions that leave no anomaly footprint. Empirical evaluations across four benchmark tabular datasets (COMPAS, German Credit, IEEE-CIS, and Communities & Crime) confirm that our method systematically crushes target feature attribution to near-zero (<0.02), maintains >90% Attack Success Rates, and fundamentally bypasses Conditional Anomaly Detection.

Comments10 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑