arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36331cs.CRcs.AIcs.LG

利用邻居校准单轮成员推断

Calibrating One-Round Membership Inference with Neighbors

  • ETH Zurich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Francesco Rita, Jie Zhang, Florian Tramèr

AI总结:

本文提出利用目标样本的邻居替代参考模型,在单轮设置下校准成员推断信号,无需额外训练,并在多个数据集上验证了其有效性与竞争力。

AI中文摘要:

最先进的成员推断(MI)方法使用参考模型(即专门训练以排除目标样本的辅助模型)对每个样本的信号进行单独校准。然而,这种范式难以扩展到现代大型模型,因为其训练成本过高而难以复制。这促使了单轮设置的出现,即仅有一个训练好的模型可用;但在没有参考模型的情况下,驱动最强攻击的逐样本校准无法再被估计,导致成员信号较弱。我们探究目标点的邻居能否在不训练任何额外模型的情况下恢复这种校准。我们的关键观察是,参考模型仅用于揭示样本在未以其训练数据训练的模型下的行为,而查询目标模型在邻近样本上的输出可获得相同的信息。我们提出了两种互补的方法来获取此类邻居,并表明针对早期训练检查点查询这些邻居可进一步锐化信号。我们在三个图像分类数据集和三种训练设置下进行了评估,结果表明邻居能产生强成员信号,且在不增加训练成本的情况下达到有竞争力的攻击性能。

英文摘要:

The state-of-the-art Membership Inference (MI) methods calibrate their signal separately for each example using reference models, auxiliary models trained to exclude the target. This paradigm scales poorly to modern large models, however, whose training is too expensive to replicate. This has motivated one-round settings, where only a single trained model is available; but without reference models the per-example calibration that drives the strongest attacks can no longer be estimated, leaving the membership signal weak. We ask whether neighbors of the target point can recover this calibration without training any additional model. Our key observation is that reference models serve only to reveal how an example behaves under models not trained on it, and that querying the target model on nearby samples yields the same information. We propose two complementary ways to obtain such neighbors, and show that querying them against an early training checkpoint further sharpens the signal. We evaluate across three image classification datasets and three training setups, showing that neighbors yield strong membership signals and competitive attack performance at no additional training cost.

↑