发表机构
Arizona State University; Williams College(亚利桑那州立大学; 威廉姆斯学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过对32个GPT-2模型的小规模预训练反事实实验,测量单示例注入对模型的短期影响,发现该影响会在训练后期衰减,且未检测到长期显著差异。
AI 中文摘要
通常我们只能估计单个训练样本对已完成模型的贡献,而非直接测量,因为测量需要运行两次成本高昂的完整预训练,仅在某一批次的某一行存在差异。我们在小规模下运行了24次这类反事实实验。我们在OpenWebText上从头开始训练了32个参数为1.24亿的GPT-2模型,分为4种条件和8个随机种子,在9536步的第200步(峰值学习率),我们将一个256行批次中的某一行替换为固定上下文注入,该注入包含一段194个token的文本。三种注入条件为:1. 具有语料库验证主语的流畅散文;2. 与上述文本主语伪造但全批次梯度差异匹配度达0.14%的流畅散文;3. 随机键盘字符。第四种条件为未注入的对照。该文本经一次曝光被模型学习后会衰减。注入后50步,8个随机种子中有8个的结果显示,见过该文本的分支对文本的交叉熵预测比未见过的分支好0.039和0.044 nat,p值均小于10^-4;在最终步,两种文本的该差异均未被检测到,p值分别为0.25和0.71,最小可检测效应为0.025和0.079 nat,且两种文本间也无差异,p值为0.54。我们报告的所有几何测量均在衰减后进行。我们预注册的插值损失屏障对比为+0.0068,p值为0.509,最小可检测效应为0.032屏障单位;保留的交叉熵为-0.00044,p值为0.310。每层中心化核对齐未在任何层检测到任何条件的可区分性。权重位移达到种子间欧氏距离的44.1%,在训练中点时已稳定92%,而屏障仅达到种子间屏障的3.0%,二者数值相差约15倍,这是下限。该注入将模型在其 basin 内重新定位,未移出。
英文摘要
A single training example's contribution to a finished model is normally estimated rather than measured, because measuring it takes two expensive full pre-training runs that differ in one row of one batch. We ran that counterfactual 24 times at a small scale. We trained 32 GPT-2 models at 124M parameters from scratch on OpenWebText, over four conditions and eight seeds. At step 200 of 9,536, at peak learning rate, we replaced one row of a 256-row batch with a fixed context injection carrying a 194-token passage. The three injected conditions are: 1. fluent prose with a corpus-attested subject, 2. fluent prose with a fabricated subject matched to it within 0.14% on full-batch gradient delta, and 3. random keyboard characters. The fourth condition is an uninjected twin. The passage is learned from one exposure and then decays. Fifty steps after injection, the arm that saw a passage predicts it better than the arm that did not by 0.039 and 0.044 nats of cross-entropy on the passage, at eight of eight seeds with p < 0.0001. At the final step we do not detect that difference for either passage, at p = 0.25 and p = 0.71, against minimum detectable effects of 0.025 and 0.079 nats, nor between the two passages, at p = 0.54. Every geometric measure we report is taken after that decay. Our pre-registered contrast on interpolation loss barrier is +0.0068 with p = 0.509, against a minimum detectable effect of 0.032 barrier units. Held-out cross-entropy is -0.00044 with p = 0.310. Per-layer centered kernel alignment does not detectably separate any condition at any layer. Weight displacement reaches 44.1% of the seed-to-seed Euclidean distance and is 92% settled by the midpoint of training, while the barrier reaches 3.0% of the seed-to-seed barrier. Those two figures sit roughly 15 times apart, and that is a lower bound. The injection relocates the model within its basin without moving it out.
Commentsv2: abstract metadata formatting fix only; paper unchanged