发表机构
Augusta University; University of Virginia; Qualcomm AI Research(奥古斯塔大学; 弗吉尼亚大学; 高通人工智能研究)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文证明冻结的Transformer可通过上下文样本模拟生成采样器,实现闭式扩散和能量基采样,并在预训练模型中观察到U形能量模式,揭示了上下文学习在数据生成中的新能力。
AI 中文摘要
越来越多的研究证明,大型语言模型不仅仅是统计记忆器,它们还具备上下文学习能力:在测试时仅利用提示中提供的示例进行推理,无需任何参数更新。先前的理论工作表明,这种能力可扩展到监督学习任务,如线性回归。我们证明,上下文学习可进一步扩展到数据生成:冻结的Transformer可以从上下文样本模拟迭代生成采样器。我们首先展示Transformer可以实现闭式和平滑闭式扩散采样器。该构造为softmax注意力确定了一个具体的生成角色:它计算责任权重和加权经验平均值,而前馈层实现Euler更新。为了将这些构造与预训练语言模型进行经验关联,我们研究了语义主题采样:提示由来自共同语义类别(如动物、食物或城市)的单词组成。在Transformer各层中,归一化隐藏状态呈现两阶段几何:在中间层向均匀球面参考移动,然后在接近输出时返回结构化的、主题相关的表示。我们进一步在这些隐藏状态云上测量相互作用粒子能量,并观察到相同的U形模式。然后我们证明Transformer可以近似基于能量的采样器,在各层中构造相同的U形能量。
英文摘要
A growing body of work establishes that large language models are not mere statistical memorizers, but are capable of in-context learning: performing inference at test time using only examples provided in the prompt, without any parameter updates. Prior theoretical work has shown that this capability extends to supervised learning tasks such as linear regression. We prove that in-context learning extends further to \emph{data generation}: frozen transformers can simulate iterative generative samplers from in-context samples. We first show that transformers can realize closed-form and smoothed closed-form diffusion samplers. The construction identifies a concrete generative role for softmax attention: it computes responsibility weights and weighted empirical averages, while feedforward layers implement Euler updates. To empirically relate these constructions to pretrained language models, we study \emph{semantic-topic sampling}: prompts consisting of words drawn from a common semantic category, such as animals, foods, or cities. Across transformer layers, the normalized hidden states exhibit a two-stage geometry: they move toward a uniform spherical reference in intermediate layers and then return to structured, topic-dependent representations near the output. We further measure an interacting-particle energy on these hidden-state clouds and observe the same U-shape pattern. We then prove that transformers can approximate an energy-based sampler, constructing the same U-shape energy across the layers.