发表机构
Karolinska Institutet(卡罗林斯卡学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对黑盒语言模型提出概率训练数据提取方法,发现聚合AUC掩盖逐文档泄露风险,发布leakit工具,强调隐私审计需按领域分解逐文档提取结果。
AI 中文摘要
语言模型的成员推断(MIA)通常用聚合ROC-AUC来总结,但这类评估存在混淆:无模型的盲基线仅通过表面文本就能将成员与非成员区分开。我们从概率视角研究黑盒、基于采样的训练数据泄露,将p(.|x)的N个样本视为输出分布的估计,并将泄露信号视为其泛函。我们将盲基线的批判扩展到采样场景:在WikiMIA上,盲词袋分类器达到AUC 0.97(5% FPR时TPR为0.90),采样未带来增益;在IID Pile拆分(MIMIR)上,自集中度或黄金延续恢复均未显著优于盲基线(增量AUC的95%置信区间包含零)。聚合指标掩盖了真实危害:相同采样能逐字提取盲攻击无法触及的尾部文档的训练数据。在Pythia-6.9B上,500个带有真实标识符的Pile文档中有83个(16.6%;带邮箱地址的文档中占21.3%)重现了该精确标识符,且在不匹配前缀的对照下未重现,因此每次泄露可归因于该文档,而非全局通用字符串。这种逐文档泄露对聚合AUC不可见,且随模型容量增长(从4.1亿参数的5.6%到69亿参数的16.6%)。风险分布不均:代码中的标识符泄露比散文强约3倍,不过散文仍明显为正且也随容量增长(从4.0%到12.1%),而任意保留延续的恢复仅局限于代码(GitHub上的成员差距为+0.44,散文上最多为+0.014)。温度和核采样影响很小,16个token的前缀就足够,且未发现语料库去重带来的减少。隐私审计应报告逐文档提取结果,按领域分解,而非单一AUC。我们发布了leakit,一款黑盒提取审计工具。
英文摘要
Membership inference (MIA) on language models is usually summarised by aggregate ROC-AUC, but such evaluations are confounded: model-free blind baselines can separate members from non-members using surface text alone. Building on probabilistic discoverable extraction, we study black-box training-data leakage using N samples from p_theta(. | x), placing mean overlap, extreme-value overlap, and self-concentration on a common functional-estimation footing. On WikiMIA, a blind bag-of-words classifier reaches AUC 0.97 (TPR 0.90 at 5% FPR) while sampling adds nothing. On an IID Pile split (MIMIR), neither self-concentration nor gold-continuation recovery significantly exceeds a blind baseline in aggregate. Aggregate metrics hide the real harm: sampling verbatim-extracts training data for a tail of documents no blind attack can reach. On Pythia-6.9B, 16.6% of 500 Pile documents bearing a real identifier (83 documents; 21.3% of those bearing an email address) have that identifier reproduced and not reproduced under a mismatched-prefix control. Each leak is attributable to that document rather than a globally common string. This per-document disclosure is invisible to aggregate AUC. Risk is uneven: identifier leakage is about 3x stronger in code than prose, though prose remains positive and grows with capacity (4.0% to 12.1% from 410M to 6.9B); recovery of arbitrary held-out continuations is essentially confined to code (+0.44 member gap on GitHub vs at most +0.014 on prose). Temperature and nucleus sampling have minor effect, a 16-token prefix suffices, and the sample-budget relationship corroborates prior probabilistic-extraction results. We detect no reduction from deduplication. Privacy audits should report per-document extraction, not only aggregate membership, and motivate differential privacy as the mitigation. We release leakit, a black-box tool implementing this probe and its control.
Comments14 pages, 7 figures. Code: https://github.com/victormaricato/leakit