高维渐近性与私有迁移学习的数据集选择
High-Dimensional Asymptotics and Dataset Selection for Private Transfer Learning
- Institute of Science and Technology Austria(奥地利科学技术研究所)
- CNRS Dauphine PSL(法国国家科学研究中心多芬纳巴黎文理研究大学)
- Chalmers University of Technology(查尔姆斯理工大学)
- University of Gothenburg(哥德堡大学)
- University of the Witwatersrand(金山大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文针对私有迁移学习中的数据集选择问题,提出基于高维回归和加权岭估计的方法,仅利用汇总统计量并保证差分隐私,通过确定性等价测试误差优化超参数,决定外部数据是否值得使用。
AI中文摘要:
要决定购买外部数据或参与协作学习,必须判断额外数据是否能充分提升预测性能以证明其成本合理。这带来了几个挑战:(i)决策通常仅依赖于公开可用的汇总统计量,而非个体层面的数据;(ii)协变量和模型偏移可能导致负迁移,使得额外数据反而降低而非提升性能;(iii)若数据敏感,其私有化需要注入噪声,这也会抵消更大样本量的收益。在本文中,我们通过多源异质高维回归和加权岭估计器对数据集选择问题进行建模。我们的方法仅使用汇总统计量,并在标签单独或特征与标签联合的情况下,以$\ ho$-零集中差分隐私提供隐私保证。主要技术贡献是测试误差的确定性等价形式,该形式捕捉了样本量、协方差结构、模型偏移、正则化和隐私噪声之间的相互作用。我们的理论允许优化超参数(权重和岭正则化器),并更广泛地决定何时私有外部数据集有用,而无需访问数据本身,仅依赖总体层面的量。这为私有迁移学习提供了理论上可处理的基础,我们通过合成和真实世界数据集的实验加以支持。
英文摘要:
To commit to buying external data or participate in collaborative learning, one must decide whether the additional data will improve prediction enough to justify the cost. This comes with several challenges: (i) the decision often relies only on aggregated statistics available publicly, rather than individual-level data; (ii) covariate and model shifts can induce negative transfer, so the additional data deteriorates rather than improves performance; (iii) if the data is sensitive, its privatization requires the injection of noise, which can also offset the benefit of a larger sample size. In this paper, we model the problem of dataset selection through high-dimensional regression with multiple heterogeneous sources and a weighted ridge estimator. Our approach uses only summary statistics and it gives privacy guarantees either on labels only or jointly on features and labels, in terms of $ρ$-zero-concentrated differential privacy. The main technical contribution is a deterministic equivalent of the test error, which captures the interactions between sample size, covariance structure, model shift, regularization and privacy noise. Our theory allows to optimize hyperparameters (weights and ridge regularizers) and, more broadly, to decide when private external datasets are useful without accessing the data itself but only relying on population-level quantities. This provides a theoretically tractable foundation for private transfer learning, which we support via experiments on both synthetic and real-world datasets.