发表机构
Dartmouth College; Breuer Lab(达特茅斯学院; 布鲁尔实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
挑战信息依赖导致训练数据隐私泄露的观点,通过三个结果揭示对抗性非鲁棒特征才是原因,引入反对抗训练建立因果关系并揭示隐私-鲁棒性权衡,修正对训练数据暴露的理解。
AI 中文摘要
在本文中,我们挑战了一种普遍观点,即信息依赖(包括死记硬背)会导致训练数据在图像重建攻击中暴露。我们表明,即使没有死记硬背,广泛的暴露也可能持续存在,相反,这是由与对抗性鲁棒性的可调连接引起的。我们首先展示了三个惊人的结果:(1)最近通过模型反演攻击(MIA)抑制重建的防御措施,在理想化攻击者下评估泄漏情况时,并没有降低信息依赖的标准度量(HSIC);(2)最大程度记住其训练数据集的模型对MIA重建仍具有鲁棒性;(3)在未看到97%训练像素的情况下训练的模型,尽管在标准假设下最近的信息理论界限给出了任意强的隐私保证,但仍可能被MIA彻底重建。为了解释这些发现,我们提供了因果证据,表明MIA下的隐私源于对抗性示例文献中所称的“非鲁棒”特征(可泛化但不可察觉且不稳定的特征)。我们进一步表明,最近的MIA防御措施通过无意中将模型转向此类特征来获得隐私改进。为了建立这种因果关系,我们引入了反对抗训练(AT-AT),这是一种训练机制,有意学习非鲁棒特征以获得比现有防御措施更好的重建防御和更高的准确性。我们的结果修正了对训练数据暴露的普遍理解,并揭示了一种新的隐私-鲁棒性权衡。
英文摘要
In this paper, we challenge the prevailing view that information dependency (including rote memorization) drives training data exposure to image reconstruction attacks. We show that extensive exposure can persist without rote memorization and is instead caused by a tunable connection to adversarial robustness. We begin by presenting three surprising results: (1) recent defenses that inhibit reconstruction by Model Inversion Attacks (MIAs), which evaluate leakage under an idealized attacker, do not reduce standard measures of information dependency (HSIC); (2) models that maximally memorize their training datasets remain robust to MIA reconstruction; and (3) models trained without seeing 97% of the training pixels, where recent information-theoretic bounds give arbitrarily strong privacy guarantees under standard assumptions, can still be devastatingly reconstructed by MIA. To explain these findings, we provide causal evidence that privacy under MIA arises from what the adversarial examples literature calls ``non-robust'' features (generalizable but imperceptible and unstable features). We further show that recent MIA defenses obtain their privacy improvements by unintentionally shifting models toward such features. To establish this causal relationship, we introduce Anti Adversarial Training (AT-AT), a training regime that intentionally learns non-robust features to obtain both superior reconstruction defense and higher accuracy than state-of-the-art defenses. Our results revise the prevailing understanding of training data exposure and reveal a new privacy-robustness tradeoff.
CommentsIn The Fourteenth International Conference on Learning Representations (ICLR'26), 2026