arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于隐私保护协同学习的带噪锚点对齐的几何数据扰动

Geometric Data Perturbation with Noisy-Anchor Alignment for Privacy-Preserving Collaborative Learning

Keiyu Nosaka, Yamato Suetake, Yuichi Takano, Yukihiko Okada, Akiko Yoshise

arXiv 2608.18749首次发表:更新:

发表机构

University of Tsukuba; Graduate School of Science and Technology, University of Tsukuba; Institute of Systems and Information Engineering, University of Tsukuba; Center for Artificial Intelligence Research, Tsukuba Institute for Advanced Research (TIAR), University of Tsukuba(筑波大学; 筑波大学大学院理工学研究科; 筑波大学系统情报工程学系; 筑波大学筑波高等研究院人工智能研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对隐私保护协同学习,提出带噪锚点对齐的几何数据扰动方法,在MNIST、CelebA实验中,其隐私-效用权衡优于对私有数据添加噪声的方案。

AI 中文摘要

几何数据扰动(GDP)支持单次隐私保护协同学习:每个参与者对其私有数据应用保距变换,仅将得到的表示上传给中央分析者。我们研究分析者与参与者合谋的GDP场景,其中分析者结合所有上传的表示、合谋参与者披露的私有数据和变换,以恢复非合谋参与者的私有数据。参与者特定的独立变换可抵御此类攻击,但会将参与者的数据映射到不兼容的表示空间,降低下游模型性能。来自数据协作(DC)分析的共享锚点对齐可恢复兼容性并提升效用,但我们表明,披露DC锚点矩阵即使在合谋存在的情况下也能精确恢复非合谋参与者的私有数据。直接对私有数据表示添加噪声可缓解此漏洞,但会大幅降低效用。我们提出改为对锚点表示添加噪声:每个参与者独立变换其私有数据和共享锚点矩阵,仅对得到的锚点表示添加噪声,并在一轮中上传两种表示。利用带噪锚点表示,分析者通过求解广义正交Procrustes问题来对齐私有数据表示。我们刻画了对齐误差和恢复误差,将保守充分条件专门化以适配我们场景下的对齐收敛,并分析了三种恢复攻击。在MNIST和CelebA上的实验表明,在评估的攻击和部署设置中,锚点噪声在可测量的泄露程度下实现了比私有数据噪声更高的学习准确率,在指定合谋模型下产生了更优的隐私-效用权衡。

英文摘要

Geometric data perturbation enables one-shot representation sharing for privacy-preserving collaborative learning: each participant applies a secret distance-preserving transformation to its private data and uploads the resulting representation to a central analyst. We study analyst-participant collusion, in which a colluding participant discloses its data and transformation to help the analyst reconstruct another participant's data. Independent participant-specific transformations block direct inversion through a disclosed common transformation but leave uploads in incompatible coordinate systems, degrading pooled learning. Data Collaboration analysis restores compatibility by aligning transformed copies of a common anchor matrix withheld from the analyst. We show that, when the centered anchor matrix has full column rank, a colluder who discloses it enables exact recovery of every participant's transformation and inversion of noiseless private representations. Adding noise to private-data representations leaves this transformation-recovery channel intact and reduces leakage at a substantial utility cost. Instead, we perturb the anchor representations: each participant perturbs only its transformed anchor representation, preserving the geometry of its private-data upload while turning known-anchor transformation recovery into a noisy estimation problem. The analyst estimates the alignment using a spectral estimator for a generalized orthogonal Procrustes problem. We analyze recovery attacks against this protocol and compare both noise placements on the CelebA and VGGFace2 facial image datasets. Under the evaluated collusion attacks, noisy-anchor alignment retains higher downstream accuracy at low identity-linkage levels. Participant-count experiments examine the utility gains and limitations of larger collaborations at comparable measured linkage.

Comments18 pages total: 13-page main paper and 5-page supplementary material

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑