QuanText:在文本数据共享中保护数据集级机密
QuanText: Protecting Dataset-Level Secrets in Textual Data Sharing
浏览论文内容
中文总结 AI 辅助
QuanText提出一种免训练、模型无关的文本数据发布机制,通过随机量化扰动全局机密及相关属性分布,在保护数据集级机密的同时保持效用,并优于现有基线。
中文摘要 AI 辅助
自然语言数据集支持许多下游应用和研究,但发布文本可能泄露底层数据源的敏感全局属性,例如与特定性别、诊断或政治立场相关的记录比例。现有工作主要集中在恢复此类全局属性的属性推断攻击上,而保护这些数据集级机密的防御措施仍然有限。差分隐私虽然能有效保护个体记录,但对聚合属性仅提供较弱的保护。我们提出了文本随机量化(QuanText),一种无需训练且与大型语言模型无关的数据发布机制,在保护文本数据集中的全局机密的同时保持数据效用。给定一个数据集级机密(如具有特定诊断的记录比例)以及需要保持效用的属性(如主题和情感),QuanText会扰动机密分布以及相关属性的分布。它通过构建关于机密和非机密属性的候选发布分布,随机选择一个与私有经验分布足够接近的候选,并使用原始文本中与属性相关的片段重写每个私有文本样本以匹配所选分布来实现。QuanText的灵感来源于统计最大泄漏(SML)框架,该框架限制了关于数据分布的机密函数的泄漏。在理想条件下,我们证明QuanText满足SML保证。由于这些条件在实践中可能不完全成立,我们还在真实数据集上对QuanText进行了实证评估。我们的结果表明,QuanText在经验隐私-效用权衡上优于竞争性的数据生成基线。
英文摘要
Natural-language datasets support many downstream applications and research studies, but releasing text can reveal sensitive global properties of the underlying data source, such as the proportion of records associated with a particular gender, diagnosis, or political stance. Existing work has largely focused on property inference attacks that recover such global properties, while defenses for protecting these dataset-level secrets remain limited. Differential privacy, although effective for protecting individual records, provides only weak protection for aggregate properties. We propose Randomized Quantization for Text (QuanText), a training-free and large-language-model-agnostic data release mechanism that protects global secrets in textual datasets while preserving data utility. Given a dataset-level secret, such as the proportion of records with a particular diagnosis, and attributes whose utility should be preserved, such as topic and sentiment, QuanText perturbs both the secret distribution and the distributions of correlated attributes. It does so by constructing candidate release distributions over secret and non-secret attributes, randomly selecting a candidate sufficiently close to the private empirical distribution, and rewriting each private text sample to match the selected distribution using attribute-related snippets from the original text. QuanText is inspired by the Statistic Maximal Leakage (SML) framework, which bounds leakage about a secret function of a data distribution. Under idealized conditions, we show that QuanText satisfies an SML guarantee. Since these conditions may not hold exactly in practice, we also evaluate QuanText empirically on real-world datasets. Our results show that QuanText achieves a better empirical privacy-utility trade-off than competing data generation baselines.
发表机构
- Carnegie Mellon University(卡内基梅隆大学)
- Microsoft Research(微软研究院)
机构由 AI 辅助整理,请以论文原文为准。