发表机构
The University of Melbourne; Macquarie University(墨尔本大学; 麦考瑞大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM成员推断基准的局限性,提出基于OLMo 2的多阶段、混杂因素受控基准OLMo-Detect,评估多种攻击,发现性能有限且受数据类型影响,对分布偏移不鲁棒。
AI 中文摘要
大型语言模型(LLM)的成员推断旨在确定给定文本样本是否包含在LLM的训练数据中,而无需访问其训练语料库。尽管最近取得了进展,现有基准存在三个局限性:训练阶段覆盖有限、成员与非成员之间的分布对齐不足,以及缺乏对非成员相对于训练语料库的严格过滤。为了解决这些局限性,我们提出了OLMo-Detect,一个基于完全开放的OLMo 2流程构建的多阶段、混杂因素受控基准。OLMo-Detect涵盖预训练、中期训练和后训练,在三个关键轴上明确对齐成员和非成员,并通过infini-gram严格过滤非成员。为了评估对分布偏移的鲁棒性,我们进一步引入了OLMo-Detect(Shifted),一种成员与非成员不对齐的变体。我们在OLMo 2系列上评估了15种无监督和3种有监督的成员推断攻击(MIA),发现:(i)整体性能有限:最佳无监督和有监督MIA的AUC均仅为0.68,且有监督MIA在跨域评估下性能下降;(ii)MIA性能在中期训练时达到峰值,在预训练和后训练时较低,这一模式由数据类型而非阶段效应驱动:精选数学数据比其他类型更易检测;(iii)总体得分从1B到13B有所提升,但在32B时趋于平稳;(iv)没有无监督MIA对分布偏移具有鲁棒性,AUC变化高达0.42。最后,我们发现我们在OLMo 2上的发现可推广到OLMo 3和非OLMo模型。
英文摘要
Membership inference on large language models (LLMs) aims to determine whether a given text sample was included in an LLM's training data, without access to its training corpus. Despite recent progress, existing benchmarks suffer from three limitations: limited coverage of training stages, insufficient distributional alignment between members and non-members, and lack of rigorous filtering of non-members against the training corpus. To address these limitations, we propose OLMo-Detect, a multi-stage, confounder-controlled benchmark built upon the fully open OLMo 2 pipeline. OLMo-Detect spans pre-training, mid-training, and post-training, explicitly aligns members and non-members on three key axes, and rigorously filters non-members via infini-gram. To assess robustness to distribution shifts, we further introduce OLMo-Detect (Shifted), a variant where members are misaligned with non-members. We evaluate 15 unsupervised and 3 supervised membership inference attacks (MIAs) across the OLMo 2 family, finding that: (i) overall performance is limited: the best unsupervised and supervised MIAs both reach an AUC of only 0.68, and supervised MIAs degrade under cross-domain evaluation; (ii) MIA performance peaks at mid-training and is lower at pre-training and post-training, a pattern driven by data type rather than a stage effect: curated math data is far more detectable than other types; (iii) overall scores improve from 1B to 13B but plateau at 32B; and (iv) no unsupervised MIA is robust to distribution shifts, with AUCs shifting by up to 0.42. Finally, we find that our findings on OLMo 2 generalize to OLMo 3 and non-OLMo models.