发表机构
University of Texas at Austin; University of Waterloo; Vector Institute(德克萨斯大学奥斯汀分校; 滑铁卢大学; 矢量研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对DP-SGD的可证明可审计隐私差距,提出轻量级防御框架提升其经验隐私且不损失理论隐私,经多场景评估验证其灵活性。
AI 中文摘要
差分隐私(DP)传统上用于为算法在训练数据发生变化时的稳定性提供理论上界。在现代隐私机器学习应用中,在效用和理论隐私之间实现良好权衡颇具挑战性,因此人们可能乐观地认为现有的理论隐私分析较为宽松。近期的隐私审计研究采用了对偶视角,转而通过构造经验区分事件来下界算法的真实隐私。迄今为止,审计领域对DP-SGD(现代机器学习中事实上的隐私训练方法)的理论隐私界的宽松性持悲观态度,因为在各种威胁模型下已实现了几乎匹配的经验下界[NHSBTJCT23, AC24, CBP25]。本研究将算法的经验隐私下界作为一个具体的优化指标,与理论上界形成互补。我们提出了一个轻量级防御框架,可通用地增强机器学习流程中的优化方法,使其在标准基准测试上的经验隐私显著提升。此外,我们证明该框架在增强DP-SGD时不会产生理论隐私成本,这与先前针对成员推断攻击提出的防御不同。我们针对广泛的审计构造、模型和数据集对防御进行评估,以证明其灵活性。
英文摘要
Differential privacy (DP) has traditionally been used to provide theoretical upper bounds on an algorithm's stability to changing its training data. In modern private machine learning applications, achieving strong tradeoffs between utility and theoretical privacy is challenging, and thus one may optimistically hope that existing theoretical privacy analyses are loose. Recent work on privacy auditing has adopted a dual viewpoint, instead lower bounding the true privacy of an algorithm by constructing empirical distinguishing events. The auditing literature has thus far yielded a pessimistic outlook on the looseness of theoretical privacy bounds for DP-SGD, the de facto private training method in modern ML, as nearly-matching empirical lower bounds have been achieved under various threat models [NHSBTJCT23, AC24, CBP25]. In this work, we propose the empirical privacy lower bound of an algorithm as a concrete metric to optimize for, complementary to the theoretical upper bound. We give a lightweight defense framework that generically augments optimization methods in the ML pipeline to have significantly-improved empirical privacy on standard benchmarks. Moreover, we show that our framework comes at no theoretical privacy cost when augmenting DP-SGD, unlike previously-proposed defenses against membership inference attacks. We evaluate our defense against a broad range of audit constructions, models, and datasets to demonstrate its flexibility.
CommentsThe code for this paper can be found here: https://github.com/pineappleEnthusiast/empirical-privacy-defense