arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

近重复族破坏精确记录成员推断

Near-Duplicate Families Break Exact-Record Membership Inference

Yiyong Liu, Jiayang Liu, Yixin Tan, Lu Sun, Rui Wen

arXiv 2609.33909首次发表:更新:

发表机构

CISPA Helmholtz Center for Information Security; Nanyang Technological University; Institute of Science Tokyo; Tohoku University(CISPA 亥姆霍兹信息安全中心; 南洋理工大学; 东京科学大学; 东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示近重复数据族使标准成员推断产生严重错误归因,提出四世界审计分离精确记录与族存在,证明正向证据仅能确认族级暴露而非精确记录来源。

AI 中文摘要

成员推断(MI)询问特定记录是否出现在模型的训练集中,并日益被用作数据来源和版权审计的证据。这些应用要求确定查询的确切记录是否被用于训练,而不仅仅是模型是否暴露于相似内容。做出这种区分具有挑战性,因为网络规模的数据集天然包含近重复内容,包括联合供稿文章、镜像页面和轻微修改的图像。我们表明,这为标准MI制造了一个根本性的混淆。干净参考审计通常针对一个零假设进行校准,在该假设中,查询记录及其近重复族均不存在。然而,在部署中,查询记录可能缺失,而一个非相同的族成员被用于训练。我们引入了一个四世界审计,独立变化精确记录包含和族存在,以区分这些情况。自然的近重复族导致严重的错误归因。在CC-News上,干净参考LiRA审计器在1.00%假阳性率下将99.70%的族存在但精确非成员标记为成员。这种失败在替代评分、模型架构以及执行的去重和重新训练中持续存在。受控干预进一步揭示该效应依赖于学习目标。在分类中,忠实族在很大程度上替代精确记录,将给定族的精确推断降至接近随机水平。在自回归语言建模中,精确序列保留可检测的残差,而族存在仍然混淆干净参考决策。这些结果表明,仅基于模型的正向成员证据可能建立族级暴露,而无法建立精确记录来源。

英文摘要

Membership inference (MI) asks whether a specific record appeared in a model's training set and is increasingly used as evidence for data provenance and copyright auditing. These applications require determining whether the exact queried record was used for training, rather than merely whether the model was exposed to similar content. Making this distinction is challenging because web-scale datasets naturally contain near-duplicates, including syndicated articles, mirrored pages, and lightly modified images. We show that this creates a fundamental confound for standard MI. A clean-reference audit typically calibrates membership against a null in which neither the queried record nor its near-duplicate family is present. In deployment, however, the queried record may be absent while a non-identical family member was used for training. We introduce a four-world audit that independently varies exact-record inclusion and family presence to separate these cases. Natural near-duplicate families cause severe false attribution. On CC-News, a clean-reference LiRA auditor labels 99.70% of family-present exact non-members as members at 1.00% false-positive rate. This failure persists across alternative scores, model architectures, and executed deduplication and retraining. Controlled interventions further reveal that the effect depends on the learning objective. In classification, faithful families largely substitute for the exact record, reducing exact-given-family inference to near chance. In autoregressive language modeling, the exact sequence retains a detectable residual, while family presence still confounds clean-reference decisions. These results show that positive model-only membership evidence may establish family-level exposure without establishing exact-record provenance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑