仅在混淆处可检测:验证重复计数揭示语言模型中的成员证据
Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models
- The University of Texas at Dallas(德克萨斯大学达拉斯分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文利用公开预训练语料库的精确重复计数,直接检验语言模型成员推断,发现普通文本中暴露痕迹微弱,高重复时与名声混淆,并揭示非成员构造中的用词偏好效应。
AI中文摘要:
当语言模型发现某个句子异常容易预测时,人们很容易得出结论认为该句子出现在其训练数据中。几乎所有已发表的针对该推断的测试都不得不猜测哪些句子在训练数据中(即成员)以及哪些不在。本文消除了这种猜测。两个模型家族,OLMo-2和Pythia,公开了其预训练语料库,并且一个针对这些语料库的公共索引能返回任何句子在每个语料库中出现的精确次数。这些计数使得三个问题可以直接回答。答案形成一把钳形攻势,从两侧合拢。在普通文本实际具有的重复水平上,从1B到13B参数的五种模型对其自身暴露仅带有微弱的痕迹。我们通过一种设计来测量该痕迹,该设计通过两个模型读取同一句子,从而在构造上抵消流畅性和质量差异,得到的秩相关接近-0.08,其中-1表示完全相关,0表示无相关。在痕迹确实变强的位置,即大约超过一千份副本时,两个语料库就哪些句子是这些句子达成一致,因为它们都是著名的句子,因此暴露无法与名声区分开来。另外两项测量展示了明显的成员信号是如何被制造的。构建非成员的一种常见方法是改变成员的一个词。模型确实偏好原始句子,但无论原始句子出现一次还是一百次,差距都是相同的,因此模型所奖励的是作者的用词选择,而非记忆。在一千份副本以上,差距随着模型规模增大而在我们可测试的十二个句子上增长,恰好在钳形攻势合拢的边界处。将对照组换成与成员在语域上不同的句子,会使检测器的AUC从0.83提高到0.94,其中0.5相当于抛硬币,1.0为完美分离。我们发布了句子库、计数和代码。
英文摘要:
When a language model finds a sentence unusually cheap to predict, it is tempting to conclude that the sentence was in its training data. Almost every published test of that inference has had to guess which sentences were in the training data, the members, and which were not. This paper removes the guessing. Two model families, OLMo-2 and Pythia, publish their pretraining corpora, and a public index over those corpora returns the exact number of times any sentence appeared in each. Those counts make three questions answerable directly. The answers form a pincer, closing from two sides. At the duplication levels ordinary text actually has, five models from 1B to 13B parameters carry at most a faint trace of their own exposure. We measure that trace with a design that reads the same sentence through two models, which cancels fluency and quality by construction, and it comes to a rank correlation near -0.08, where -1 would be a perfect relation and 0 none. Where the trace does become strong, above roughly a thousand copies, the two corpora agree on which sentences those are, because they are the famous ones, so exposure can no longer be told apart from fame. Two further measurements show how apparent membership signal gets manufactured. A common way to build a non-member is to change one word of a member. The model does prefer the original, but the gap is the same, within noise, whether the original appeared once or a hundred times, so what the model is rewarding is the author's word choice, not memory. Above a thousand copies the gap grows with model size on the twelve sentences we can test there, at the same boundary where the pincer closes. And swapping the controls for sentences that differ from the members in register moves a detector from 0.83 to 0.94 AUC, on a scale where 0.5 is a coin flip and 1.0 is perfect separation. We release the sentence banks, counts, and code.