发表机构
University of Chicago(芝加哥大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过分析Pythia等模型在间接宾语识别任务中的早期训练窗口,发现重复名字偏好源于注意力头计算的暂时失衡,并证明移植成熟参数可部分恢复性能,揭示成熟电路掩盖早期因果配置的现象。
AI 中文摘要
机制可解释性通常研究完全训练好的模型,然而,驱动某种行为的计算在模型仍在学习任务时可能会发生变化。在间接宾语识别任务中,模型应继续使用仅提及一次的名字,而非提及两次的名字。Pythia 模型会经历一个早期训练窗口,在此期间它们偏好重复的名字,因此在两个名字之间进行选择的准确率会降至二分之一以下,而同一窗口内固定文本样本上的语言模型损失却持续下降。该窗口反映了两种计算之间的暂时性不平衡。我们通过选择一组在独立于因果评估所用提示的提示上写入重复名字的注意力头,保持该选择固定,然后将每个头的最终令牌输出替换为其在另一组非重复名字提示上的平均输出,从而识别出错误偏好的一种原因。这改善了在单独训练的 160M 模型以及官方 160M、410M 和 1B 模型中的正确减去重复的 logit 差值。在 160M 规模下,成熟模型中降低重复名字的头在此阶段几乎未表现出其成熟行为。它仅将不到百分之一的注意力分配给重复提及,其输出对降低该名字的 logit 几乎没有直接贡献。这两个属性在行为恢复的时间间隔内均有所增长。在 160M、410M 和 1B 模型中,将相应头的成熟参数移植到早期检查点可恢复训练结束时观察到的正确减去重复 logit 差值总改善的 35% 至 68%。相关的早期到晚期反转出现在更多 Pythia 规模、两个独立训练的 GPT-2 模型以及 OLMo 中。因此,一个成熟的电路可能掩盖了早期训练中塑造行为的瞬态因果配置。
英文摘要
Mechanistic interpretability usually studies fully trained models, yet the computations that drive a behaviour can change while the model is still learning the task. On the Indirect Object Identification task, a model should continue with the name mentioned once rather than the name mentioned twice. Pythia models pass through an early training window in which they prefer the repeated name, so accuracy in a choice between the two names falls below one half while language-model loss on a fixed text sample keeps decreasing across the same window. The window reflects a temporary imbalance between two computations. We identify one cause of the wrong preference by selecting a set of attention heads that write the repeated name, on prompts separate from those used for causal evaluation, keeping that selection fixed, and then replacing each head's final-token output with its average output on a separate set of non-repeated-name prompts. This improves the correct-minus-repeated logit difference in a separately trained 160M model and in the official 160M, 410M, and 1B models. At 160M, the head that lowers the repeated name in the mature model shows little of its mature behaviour at this point. It directs less than one percent of its attention to the repeated mention, and its output makes almost no direct contribution to lowering that name's logit. Both properties grow over the interval in which behaviour recovers. Across the 160M, 410M, and 1B models, transplanting the corresponding head's mature parameters into the early checkpoint recovers 35 to 68 percent of the total improvement in the correct-minus-repeated logit difference seen by the end of training. Related early-to-late reversals appear at further Pythia scales, in two independently trained GPT-2 models, and in OLMo. A mature circuit can therefore conceal a transient causal configuration that shaped behaviour earlier in training.
CommentsAccepted to Findings of EMNLP 2026