arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

立场:隐私是一种声明,而非合成数据的属性

Position: Privacy Is a Claim, Not a Property of Synthetic Data

Jiachen Zhao, Antonia Januszewicz, Taeho Jung

arXiv 2609.01273首次发表:更新:

发表机构

University of Notre Dame(圣母大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对合成数据隐私属性的认知偏差,经实证分析发现其隐私保证隐含难验证,主张将隐私视为明确的证据型科学声明,建议ML venues制定相关规范。

AI 中文摘要

合成数据已成为机器学习研究的常见组成部分。尽管被广泛采用,但其在隐私敏感场景中的应用已悄然从“在既定假设下存在残余推理风险”的声明,转变为“从数据生成本身推断出的基于外观的属性”。在这篇立场论文中,我们认为这种转变反映了社区标准中“何为充分隐私证据”的隐含变化,而非对成熟隐私原则的误解。通过对近期主要机器学习会议论文的实证分析,我们发现合成数据常被用于隐私敏感场景,却未明确阐明威胁模型、推理风险或可证伪的隐私声明。因此,隐私保证往往是隐含的,难以验证且分布不均,稀有和少数记录的暴露风险更高。我们主张将隐私视为明确、基于证据的科学声明,并建议机器学习会议采用规范,要求与隐私相关的断言必须范围清晰、可测试且可争议。

英文摘要

Synthetic data has become a common component of machine learning research. While widely adopted, its use in privacy-sensitive contexts has quietly shifted from a claim of residual inference risk under stated assumptions to an appearance-based property inferred from data generation itself. In this position paper, we argue that this shift reflects an implicit change in community standards for what counts as sufficient privacy evidence, rather than a misunderstanding of well-established privacy principles. Drawing on an empirical analysis of recent publications across major ML venues, we show that synthetic data is frequently used in privacy-sensitive settings without explicit articulation of threat models, inference risks, or falsifiable privacy claims. As a result, privacy assurance often remains implicit, difficult to verify, and unevenly distributed, with heightened exposure for rare and minority records. We argue for treating privacy as an explicit, evidence-based scientific claim and recommend that ML venues adopt norms requiring privacy-relevant assertions to be clearly scoped, testable, and contestable.

Journal refProceedings of the 43rd International Conference on Machine Learning (ICML 2026), 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑