arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27037cs.CR

邻域监测:种子局部组合合成数据中的隐私风险

Neighborhood Watch: Privacy Risks in Seeded Local Combination Synthetic Data

Hadrien Lautraite, Tristan Allard, Anne-Sophie Charest, Jean-François Rajotte, Sébastien Gambs

AI总结:

本文针对医疗领域用于共享“匿名化数据”的SMOTE、Simulant和Avatar三种种子局部组合合成数据生成方法,开展多种隐私攻击分析,发现其存在大量隐私泄露,质疑其输出的匿名性。

AI中文摘要:

合成数据被视为敏感场景下数据共享的有前景解决方案,但近期隐私攻击研究表明其仍存在显著残余风险,尤其是未基于差分隐私等形式化方法的合成数据生成方法。本文研究种子局部组合合成数据生成方法的隐私风险,该方法通过组合真实邻域样本构建合成样本,具体聚焦该类中的SMOTE、Simulant和Avatar三种方法,它们近期被用于医疗领域共享“匿名化数据”。本文通过成员推理、链接和重构等多种攻击开展广泛隐私分析,结果显示三种方法均存在大量隐私泄露,引发对其输出在实践中是否应被视为匿名的严重质疑。

英文摘要:

Synthetic data is seen as a promising solution for sharing data in sensitive contexts. However, recent work on privacy attacks have shown that there are still significant residual risks, especially for synthetic data generations methods that are not based on formal approaches such as differential privacy. In this paper, we investigate the privacy risks associated with local combination approaches for generating synthetic data in which synthetic profiles are built by combining real neighbouring profiles. More precisely, we focus on three methods from this family, namely SMOTE, Simulant and Avatar, which have been recently used as a way to share 'anonymised data' in the healthcare domain. In particular, we conduct an extensive privacy analysis through a diverse set of attacks: membership inference, linkage and reconstruction attacks. Our results demonstrate substantial privacy leakage for all three methods, raising serious doubts about whether their outputs should be regarded as anonymous in practice.

补充信息

↑