发表机构
Boston University(波士顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出注意力均值场分析,用平均注意力核预测表示几何演化,并通过均值场偏差揭示上下文特定计算,在归纳和少样本任务中验证了其有效性。
AI 中文摘要
语言模型的表示几何并非预先确定;它随着模型运行而演化。对该几何的忠实描述必须捕捉这一动态过程,因此不能仅基于与模型无关的统计量(如共现)。在此,我们引入对注意力的均值场分析。从一个词元到另一个词元的平均注意力定义了一个核,该核将表示逐层传递,并可在网络中迭代以建模几何如何被变换。我们以两种方式对该平均值进行条件化。当以整个语料库为条件时,该核预测表示几何的平均情况演化。当以单个上下文为条件时,它预测该上下文的期望几何。一个注意力头偏离该预测的程度,即其“均值场偏差”,隔离了均值场所遗漏的上下文特定计算。在语料库条件解读下,该核产生一个开环模型:仅从输入嵌入和冻结权重出发,我们即可在词元表示上迭代该核和模型自身的多层感知机,而无需在任何层查阅实测偏差。所得预测高度准确。在训练早期,模型与其语料库均值场不可区分。将每个注意力头替换为其均值场,该替换在真实文本上不改变损失。在归纳(induction)出现前后,两者开始分化,且随着表示变得上下文化,差距扩大。在上下文条件解读下,偏离均值场是上下文特定计算的任务无关度量。残差可加性分解为异常注意力路由和传输值的上下文化。在受控归纳和少样本设置中,更大的偏差对应着对上下文信息更强的依赖。
英文摘要
A language model's representation geometry is not predetermined; it evolves as the model runs. A faithful account of that geometry must capture that dynamic process, and so cannot be based solely on model-independent statistics such as co-occurrence. Here we introduce a mean-field analysis of attention. The average attention from one token to another defines a kernel that carries representations layer to layer and can be iterated through the network to model how the geometry is transformed. We condition this average two ways. Conditioned on a whole corpus, the kernel predicts the average-case evolution of representation geometry. Conditioned instead on a single context, it predicts the expected geometry for that context. A head's departure from that prediction, its \emph{mean-field deviation}, isolates the context-specific computation that the mean field misses. Under the corpus-conditional reading, the kernel yields an open-loop model: from the input embeddings and the frozen weights alone, we can iterate the kernel and the model's own MLPs over token representations, never consulting a measured deviation at any layer. The resulting prediction is highly accurate. In early training the model and its corpus mean field are indistinguishable. Replace every attention head with its mean field, and the substitution leaves the loss on real text unchanged. Around the onset of induction, the two diverge, and the gap widens as representations become contextualized. Under the context-conditional reading, deviation from the mean field is a task-agnostic measure of context-specific computation. The residual decomposes additively into unusual attention routing and contextualization of the transported values. Across controlled induction and few-shot settings, greater deviation tracks greater reliance on in-context information.