arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10627cs.CRcs.LG

SoK:通过可解释人工智能对机器学习进行隐私攻击

SoK: Privacy Attacks on Machine Learning via Explainable AI

Abdullah Caglar Oksuz, Anisa Halimi, Erman Ayday

首次发表
浏览论文内容

中文总结 AI 辅助

本研究系统化25项利用可解释AI信号(如解释统计、反事实距离)进行模型提取、成员推断和模型反演的攻击,提出五条获取路径,并指出风险取决于信号暴露方式,主张按端到端披露问题评估隐私防御。

中文摘要 AI 辅助

机器学习解释揭示了预测之外模型行为,为模型机密性和数据隐私创造了攻击面。我们系统化了25项利用解释进行模型提取、成员推断和模型反演的研究,将属性推断视为部分反演。现有工作通常仅标记为黑盒或白盒,掩盖了对手获得的解释信号在实质上的差异。因此,我们将模型知识与解释获取分离,并识别出五条路径:目标发布、攻击者推导、二次披露、特权访问和发布的全局工件。在这些路径中,解释降低了提取成本,通过解释统计、反事实距离和解释引导的鲁棒性暴露成员信号,并支持对私有输入的空间或代数重建。我们比较了系统和威胁模型、解释信号、辅助知识、目标模型、模态、查询预算、评估指标、报告性能和防御措施。我们的分析表明,没有一种解释族是普遍不安全的,也没有一种防御是普遍有效的。风险取决于暴露了哪种信号、如何获取、针对哪种资产以及攻击者已知的信息。我们认为,解释隐私因此应被评估为一个端到端的披露问题,防御措施应与获取路径和受保护资产相匹配。

英文摘要

Machine learning explanations reveal model behavior beyond predictions, creating attack surfaces for model confidentiality and data privacy. We systematize 25 studies that exploit explanations for model extraction, membership inference, and model inversion, treating attribute inference as partial inversion. Existing work is often labeled only black- or white-box, obscuring substantial differences in what explanation signal reaches an adversary. We therefore separate model knowledge from explanation acquisition and identify five paths: target-released, attacker-derived, secondary disclosure, privileged access, and released global artifacts. Across these paths, explanations reduce extraction cost, expose membership signals through explanation statistics, recourse distance, and explanation-guided robustness, and support spatial or algebraic reconstruction of private inputs. We compare system and threat models, explanation signals, auxiliary knowledge, target models, modalities, query budgets, evaluation metrics, reported performance, and defenses. Our analysis shows that no explanation family is uniformly unsafe and no defense is uniformly effective. Risk depends on which signal is exposed, how it is acquired, which asset is targeted, and what the attacker already knows. We argue that explanation privacy should therefore be evaluated as an end-to-end disclosure problem, with defenses matched to the acquisition path and protected asset.

发表机构

  • Case Western Reserve University(凯斯西储大学)
  • IBM Research(IBM研究院)

机构由 AI 辅助整理,请以论文原文为准。

↑