arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14306cs.AI

利用经验下一个token分布追踪大语言模型行为到训练数据

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions

Zachary Izzo

首次发表
浏览论文内容

中文总结 AI 辅助

研究大语言模型输出分布与训练数据的联系,通过经验下一个token分布追踪。发现多数输入下二者近乎完美一致且随模型规模等提升,也存在差异,探讨了差异来源,鼓励以数据为中心的机制可解释性研究。

中文摘要 AI 辅助

在本文中,我们研究了大语言模型(LLM)的输出分布与其训练数据之间的联系。具体而言,我们研究了在给定训练数据中的上下文时,LLM的下一个token分布与经验下一个token分布(ENTD)的一致程度。ENTD是一个有吸引力的目标,因为它是预训练中使用的下一个token交叉熵损失的无限制全局最小值,也是预训练语料库的一个易于解释的函数。我们发现,对于很大一部分输入,LLM的分布与ENTD几乎完美一致,并且平均一致性随着模型规模和训练计算量的增加而提高。然而,存在一长串输入序列,其中LLM和ENTD有显著差异,我们研究了这种差异在Transformer架构、训练过程以及ENTD估计本身的有限样本噪声中的几个可能来源。更广泛地说,我们希望我们的发现将鼓励更多关于“以数据为中心的机制可解释性”的工作,这是对标准机制可解释性的补充,它打开了模型行为如何从数据中产生而不是如何在学习权重中编码的黑箱。

英文摘要

In this paper, we study the connection between an LLM's output distribution and the data used to train it. Specifically, we study the degree to which an LLM's next-token distribution agrees with the empirical next-token distribution (ENTD) given the context in the training data. The ENTD is an appealing target because it is the unrestricted global minimizer of the next-token cross entropy loss used for pretraining, as well as an easily interpretable function of the pretraining corpus. We find that for a significant fraction of inputs, the LLM's distribution agrees with the ENTD almost perfectly, and the agreement generally increases with model scale and training compute. Nevertheless, there is a long tail of input sequences where the LLM and ENTD differ significantly, and we examine several possible sources of this discrepancy across the transformer architecture, training procedure, and finite-sample noise in the ENTD estimate itself. More broadly, we hope our findings will encourage more work on ``data-centric mechanistic interpretability,'' a complement to standard mechanistic interpretability that opens the black box of how model behaviors arise from the data, rather than how they are encoded in the learned weights.

发表机构

  • NEC Labs America(美国 NEC 实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑