公共数据集的差分隐私混合提升私有学习
Differentially Private Mixing of Public Datasets Improves Private Learning
浏览论文内容
中文总结 AI 辅助
提出首个私有学习多个公共数据集混合以预训练的流水线,通过低维线性模型实现,在NIH和ENRON数据集上显著提升效用。
中文摘要 AI 辅助
许多机器学习应用涉及敏感数据,因此需要在差分隐私(DP)下进行训练。然而,DP训练常常会降低模型效用。在某些情况下,先在“公共”数据上预训练模型,再在敏感数据上使用DP进行微调,可以减少效用的下降。然而,这取决于所选公共数据集与敏感数据的相关性。我们引入了第一个流水线,该流水线针对给定的敏感下游任务,私有地学习多个公共数据集的混合以进行预训练。我们的关键见解是,我们可以通过私有地学习一个低维线性模型,来私有地找到多个公共数据集的最佳混合。我们在NIH数据集上测试了我们的方法,用于X射线分类,以及在ENRON电子邮件数据集上用于语言建模。应用我们的方法为NIH ChestX-ray14数据集中的疾病找到定制的X射线数据集混合进行预训练,与基线相比,在各种隐私预算下,我们将宏AUC提高了最多0.037,在Cardiomegaly上,在ε=1时,相对AUC增益高达+22.8%。对于在ENRON数据集上的DP训练,在我们的混合数据集(The Common Pile,一个公共领域文本数据集的集合)上进行预训练,相对于基线混合,测试困惑度降低了16%。
英文摘要
Many machine learning applications involve sensitive data and therefore require training under differential privacy (DP). However, DP training often degrades model utility. In some cases, first pre-training the model on "public" data before finetuning with DP on the sensitive data can reduce the drop in utility. However, the success of this depends on how relevant the selected public dataset is to the sensitive data. We introduce the first pipeline that privately learns the mixture of several public datasets to pretrain on for a given sensitive downstream task. Our key insight is that we can privately find the best mixture of multiple public datasets by privately learning a low-dimensional linear model. We tested our method on the NIH dataset for X-ray classification and the ENRON email dataset for language modeling. Applying our method to find tailored mixtures of X-ray datasets to pretrain on for diseases in the NIH ChestX-ray14 dataset, we improved macro AUC by up to 0.037 across privacy budgets compared to the baselines, with gains as large as +22.8% relative AUC on Cardiomegaly at $ε=1$. For DP training on the ENRON dataset, pre-training on our mixture of The Common Pile (a collection of public-domain text datasets) decreased test perplexity by 16% relative to the baseline mixtures.
发表机构
- University of Toronto(多伦多大学)
- Vector Institute(向量研究所)
- CISPA Helmholtz Center for Information Security(CISPA亥姆霍兹信息安全中心)
机构由 AI 辅助整理,请以论文原文为准。