arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当文本到图像合成数据适得其反时:真实-合成混合训练中的隐私风险放大

When T2I Synthetic Data Backfires: Amplified Privacy Risks in Real-Synthetic Mix Training

Na Li, Boyu Kuang, Hongsheng Hu, Liquan Chen, Hyoungshick Kim, Yansong Gao, Anmin Fu

arXiv 2607.13541首次发表:更新:

发表机构

Nanjing University of Science and Technology; The University of Newcastle; Southeast University; Sungkyunkwan University; The University of Western Australia(南京理工大学; 纽卡斯尔大学; 东南大学; 成均馆大学; 西澳大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究真实-合成混合训练中合成数据会放大真实训练样本隐私泄露问题,建立理论框架,提出RSMixLeak通过成员推理攻击评估风险,有两个变体,还给出轻量级泄露倾向指标以识别高风险数据集。

AI 中文摘要

为克服数据收集的数据稀缺和隐私限制,学术界和工业界的标准做法是用文本到图像(T2I)生成的合成数据增强真实训练数据,即真实-合成混合训练(RSMT)范式。虽然用合成数据替代敏感真实样本被广泛视为减轻被替代数据隐私暴露的手段,但对积极参与训练的其余真实样本的风险基本未被研究。本文首次揭示RSMT会大幅放大这些真实训练样本的隐私泄露。建立理论框架RSMT记忆放大,证明纳入合成数据会使真实样本向混合特征空间的外围区域移动,迫使模型更积极地记忆它们。在此基础上,提出RSMixLeak通过成员推理攻击(MIAs)系统评估此风险。RSMixLeak有两个变体,非对抗性变体用诚实的T2I提供者审计良性RSMT管道,建立真实数据与T2I生成数据之间内在差距引起的泄露下限。对抗性变体考虑控制T2I模型或向T2I提供者贡献精心制作数据的对手,通过高级语义属性绑定或不可察觉的像素级涂层故意扩大目标类上的分布差距,进一步放大真实训练数据上的泄露同时提高下游模型效用。基于这些发现,还提出仅从真实数据可计算的轻量级泄露倾向指标,可靠识别不适用于进入RSMT的高风险数据集,作为自我评估的缓解措施。

英文摘要

To overcome data scarcity and privacy constraints in data collection, it has become standard practice across academia and industry to augment real training data with text-to-image (T2I)-generated synthetic data, a paradigm we term Real-Synthetic Mix-Training (RSMT). While substituting synthetic data for sensitive real samples is widely regarded as a means to mitigate privacy exposure of the substituted data, the risk to the remaining real samples that actively participate in training has remained largely unexamined. This work reveals, for the first time, that RSMT can substantially amplify privacy leakage of these real training samples. We establish a theoretical framework, RSMT Memorization Amplification, proving that incorporating synthetic data displaces real samples toward peripheral regions of the mixed feature space, in turn forcing the model to memorize them more aggressively. Guided by this foundation, we propose RSMixLeak to systematically assess this risk through membership inference attacks (MIAs). RSMixLeak comprises two variants depending on the adversary's capability. The non-adversarial variant audits a benign RSMT pipeline with an honest T2I provider, establishing a lower bound on the leakage induced by the intrinsic gap between real and T2I-generated data. The adversarial variant considers an adversary who controls the T2I model or contributes crafted data to the T2I provider, and deliberately enlarges this distributional gap on a target class via either high-level semantic attribute binding or imperceptible pixel-level coating, further amplifying leakage on real training data while improving downstream model utility. Motivated by these findings, we further propose a lightweight leakage propensity indicator computable from real data alone that reliably identifies high-risk datasets unsuitable for entering RSMT, as a self-assessable mitigation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑