发表机构
Yale; AVERI; Stanford; Google Research; Google DeepMind; Cornell(耶鲁大学; AVERI(未提及常见中文名,可直译为“应用价值与伦理研究所”之类,具体需结合实际背景); 斯坦福大学; 谷歌研究院; 谷歌DeepMind; 康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型可提取记忆的有效性问题,提出通过匹配比较测量训练与非训练序列生成概率来判断记忆,形式化共形测试和普查两种匹配比较方式并揭示问题,完善可提取记忆定义,给出有效声明及现实预算下的生成要求。
AI 中文摘要
近期关于大语言模型中可提取记忆的研究存在两个相互矛盾的有效性问题。一些研究夸大了提取效果,比如依赖过短序列难以区分记忆和可预测性。另一些则暗示提取不是记忆的可靠证据,因为模型也能再现未明确训练过的真实世界文本。两种情况都忽略了有效提取声明的关键:模型必须以足够高的概率生成训练序列以表明记忆。为此需进行匹配比较,测量感兴趣的训练序列和可比非训练序列的生成概率。非训练序列概率提供可预测性基线,超过此基线的训练序列即为记忆证据。我们通过两种方式形式化匹配比较:一是共形测试,从总体中采样训练和非训练序列时校准阈值到选定的误报率;二是普查,针对单个文档(如一本书)校准到匹配的非训练文档。我们表明匹配比较能实现严格、校准的记忆声明,并揭示先前设置的有效性问题。例如,在维基百科上,OLMo 2 32B再现非训练10词后缀的频率约为训练后缀的24%,这反映的是误报而非记忆。对于Llama 3.1 70B在书籍上的情况,我们校准的阈值低至1e - 27,支持了在实际采样预算下无法提取的序列的记忆声明。基于这些结果,我们完善了“可提取记忆”的定义,要求有有效的记忆声明且在现实预算内几乎确定的生成。
英文摘要
Recent work on extractable memorization in language models suffers from two contrasting validity problems. Some studies overstate extraction, for example, by using sequences too short to distinguish memorization from predictability. Others imply that extraction is unreliable evidence because models can reproduce real-world text they weren't explicitly trained on. Both overlook what makes a valid extraction claim: the model must generate a training sequence with high enough probability to indicate memorization. To determine what's high enough, one has to perform a matched comparison, measuring generation probabilities of training and comparable non-training sequences. Since non-training sequences can't have been memorized, their probabilities are a baseline for predictability; exceeding this baseline is evidence of memorization. We formalize matched comparisons with (1) a conformal test calibrated to a chosen false-positive rate when sequences are sampled from populations, and (2) a single-document census that calibrates against a matched non-training document. Matched comparisons enable rigorous, calibrated memorization claims and clarify where prior setups have validity issues. On Wikipedia, OLMo 2 32B reproduces non-training 10-token suffixes roughly 24% as often as training ones: that share reflects false positives, not memorization. For Llama 3.1 70B on books, calibrated census thresholds reach as low as 10^(-27), supporting memorization claims for sequences no feasible sampling budget would extract. We therefore refine "extractable memorization" to require both a valid memorization claim and near-certain generation within a realistic budget.