arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10611cs.CRcs.LGstat.ML

通过词级概率估计的黑盒成员推断

Black-Box Membership Inference via Word-Level Probability Estimation

Shengjie Niu, Yeheng Ge, Jian Huang

首次发表
浏览论文内容

中文总结 AI 辅助

针对仅暴露文本续写的专有大模型,提出基于词级概率估计的黑盒成员推断方法WPMIA,通过蒙特卡洛采样和前缀条件似然放大成员差异,在开源和专有模型上优于现有基线。

中文摘要 AI 辅助

成员推断攻击(MIAs)已成为审计大型语言模型(LLMs)隐私风险的关键工具,旨在确定给定文本是否包含在模型的训练语料库中。然而,大多数现有的MIAs需要访问逐词元的对数几率或概率,这使得它们在实际中不适用于仅暴露文本续写的专有LLMs。为解决这一尚未充分探索的设置,我们提出了词级概率MIA(WPMIA),一种用于严格黑盒隐私审计的统计上有原则的MIA。WPMIA通过蒙特卡洛采样和局部核平滑估计词级生成概率,然后将这些估计聚合为序列级似然估计器。此外,WPMIA构建了在不同前缀条件下的似然,从而放大了成员与非成员之间的分布差异。我们在各种开源LLMs上评估了WPMIA,发现它始终优于现有的黑盒基线。重要的是,我们还在现代专有LLMs上评估了WPMIA,包括GPT-5-Chat、Gemini-2.5-Flash和Claude-4.5-Haiku,在这些模型上实现了平均TPR@5%FPR为42.0。这些结果为未来严格黑盒成员推断的研究提供了坚实的基础。代码可在 \nhref{ this https URL }{ this https URL } 获取。

英文摘要

Membership inference attacks (MIAs) have emerged as critical tools for auditing privacy risks in large language models (LLMs), aiming to determine whether a given text was included in a model's training corpus. However, most existing MIAs require access to per-token logits or probabilities, making them inapplicable in practice to proprietary LLMs that expose only textual continuations. To address this underexplored setting, we propose Word-level Probability MIA (WPMIA), a statistically principled MIA for strict black-box privacy auditing. WPMIA estimates word-level generation probabilities via Monte Carlo sampling with local kernel smoothing, then aggregates these estimates into a sequence-level likelihood estimator. Furthermore, WPMIA constructs the likelihood conditioned on different prefixes, thereby amplifying the distributional differences between members and non-members. We evaluate WPMIA across various open-source LLMs and find that it consistently outperforms existing black-box baselines. Importantly, we also evaluate WPMIA on modern proprietary LLMs, including GPT-5-Chat, Gemini-2.5-Flash, and Claude-4.5-Haiku, achieving an average TPR@5\%FPR of 42.0 across these models. These results offer a sound foundation for future research on strict black-box membership inference. Code is available at \href{https://github.com/niusj03/WPMIA}{https://github.com/niusj03/WPMIA}.

发表机构

  • The Hong Kong Polytechnic University(香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑