我见过你?嵌入行为信号的合成人脸数据集成员推断
Have I Seen You? Embedding Behavior Signals Synthetic Face Dataset Membership
- University of Luxembourg(卢森堡大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对合成人脸数据集开展数据集级成员推断攻击,在多模型多数据集实验中实现高成功率,揭示合成数据仍保留真实训练数据痕迹,需更强隐私泄漏缓解措施。
AI中文摘要:
合成人脸数据集越来越多地用于减少生物识别中的隐私暴露和数据访问限制。然而,生成这些数据集的模型是在真实人脸数据上训练的,因此合成数据仍可能泄露其真实源数据。我们通过一种数据集级成员推断攻击研究这一风险,该攻击首先识别用于训练人脸识别模型的合成数据集,再推断生成器所用的真实数据集。在11个人脸识别模型、11个合成数据集和7个真实数据集的实验中,该攻击在100%的案例中成功恢复了合成训练数据集,在54.5%的案例中识别出生成器的源数据集。这些结果表明,合成数据可保留真实训练数据的数据集级痕迹,隐私保护部署需更强的泄漏缓解措施。
英文摘要:
Synthetic face datasets are increasingly used to reduce privacy exposure and data access constraints in biometric recognition. Yet the generators that produce these datasets are trained on real faces, so synthetic data may still reveal their real source data. We study this risk through a dataset-level membership inference attack that first identifies the synthetic dataset used to train a face recognizer and then infers the real dataset used to train the generator. Across 11 face recognition models, 11 synthetic datasets, and 7 real datasets, the attack recovers the synthetic training dataset in 100% of cases and identifies the generator's source dataset in 54.5% of cases. These results show that synthetic data can retain dataset-level traces of real training data and that privacy-preserving deployment requires stronger leakage mitigation.