DIME:针对扩散模型的查询高效成员推断框架
DIME: Query-Efficient Framework for Membership Inference on Diffusion Models
- University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出理论依据充分、查询高效的DIME框架,可针对扩散模型开展成员推断攻击,在多数据集上的性能优于现有攻击,还提出并评估了对应防御措施。
AI中文摘要:
成员推断攻击可揭示个体记录是否被用于训练模型,但现有针对扩散模型的攻击大多是启发式的,且需要大量查询预算。我们提出DIME(去噪器理想成员误差,Denoiser Ideal Membership Error),这是一种有理论依据且查询高效的扩散模型成员推断框架。我们的研究起点是对有限训练集的最优扩散去噪器的精确刻画,该刻画表明成员泄漏由去噪器的隐式重构误差决定。此误差可分解为两个互补信号:捕捉重构精度的偏置项,以及此前未被探索的、捕捉附近训练样本几何结构的局部拥挤项。两者均可仅通过模型查询实现高效估计,从而得到仅需两次查询的实用攻击。在CIFAR-10/100、STL10-U、CelebA和ImageNet数据集上,DIME在相当或显著更低的查询成本下始终优于现有攻击,在1% FPR(假阳性率)下的TPR(真阳性率)提升了最多3倍;值得注意的是,其两次查询变体可优于现有需30次查询的基线。最后,我们提出、讨论并评估了可抵御此类强大成员测试的特定防御措施。
英文摘要:
Membership inference attacks expose whether individual records were used to train a model, yet existing attacks on diffusion models are largely heuristic and can require substantial query budgets. We introduce DIME (Denoiser Ideal Membership Error), a theoretically grounded and query-efficient framework for membership inference on diffusion models. Our starting point is an exact characterization of the optimal diffusion denoiser for a finite training set, which reveals that membership leakage is governed by the denoiser's implicit reconstruction error. This error decomposes into two complementary signals: a bias term, capturing reconstruction accuracy, and a previously unexplored local crowding term, capturing the geometry of nearby training examples. Both admit efficient estimators using only model queries, yielding a practical attack with as few as two queries. Across CIFAR-10/100, STL10-U, CelebA, and ImageNet, DIME consistently outperforms prior attacks at comparable or substantially lower query cost, improving TPR at 1% FPR by up to $3\times$; remarkably, its two-query variant can outperform existing 30-query baselines. Finally, we suggest, discuss, and evaluate specific defenses to counteract such powerful membership tests.