上下文阶梯:小初始化下词元嵌入的签名对齐动力学
Context Staircase: Signature-Aligned Dynamics of Token Embeddings under Small Initialization
浏览论文内容
中文总结 AI 辅助
本研究提出上下文阶梯现象,发现小初始化下词元嵌入从低阶到高阶逐步对齐概率签名,推导相关演化方程并揭示数据统计与架构对嵌入的塑造作用。
中文摘要 AI 辅助
词元嵌入是连接语言模型中离散词元与连续计算的基本表示单元。尽管现代语言模型通过基于梯度的训练从随机初始化中学习嵌入,但有意义的嵌入结构如何产生的动力学机制仍不清楚。本研究中,我们发现演化的嵌入结构与词元条件标签和上下文分布密切相关,我们将其形式化为概率签名。我们观察到一种渐进学习过程,将其命名为上下文阶梯:嵌入先学习数据的低阶统计签名,再学习高阶统计签名。更具体地说,我们观察到训练早期,嵌入与将词元与其标签关联的最简单无上下文签名对齐;随着训练推进,它们逐步反映涉及越来越多上下文词元的签名。随后,我们分析小初始化下嵌入的梯度流,推导了前馈和自注意力架构的嵌入演化方程。我们还将这些观察扩展到实际语言模型训练中。最后,我们证明这些嵌入结构在任务学习和将语义结构融入嵌入空间两方面发挥重要作用。总体而言,我们的结果为数据统计与架构如何共同塑造语言模型中的词元嵌入提供了动态解释,并揭示了数据统计空间中的一种隐含偏置:训练从更简单的低阶统计关系向日益复杂的上下文依赖关系推进。
英文摘要
Token embeddings are the basic representational units that connect discrete tokens with continuous computation in language models. Although modern language models learn embeddings from random initialization through gradient-based training, the dynamical mechanism by which meaningful embedding structures emerge remains unclear. In this work, we identify that the evolving embedding structures are closely related to token-conditioned label and contextual distributions, which we formalize as probability signatures. We observe a progressive learning process, which we term Context Staircase: embeddings learn the low-order statistic signatures of the data before the high-order ones. More specifically, we observe that early in training they align with the simplest, context-free signature linking a token to its label, and as training proceeds, they progressively reflect signatures involving more and more context tokens. We then analyze the gradient flow of embeddings under small initialization to explain this phenomenon, deriving embedding evolution equations for feed-forward and self-attention architectures. We further extend these observations to real language-model training. Finally, we show that these embedding structures play an important role in both task learning and the incorporation of semantic structure into the embedding space. Overall, our results provide a dynamic explanation of how data statistics and architecture jointly shape token embeddings in language models, and reveal an implicit bias in the space of data statistics: training proceeds from simpler, low-order statistical relations toward increasingly complex, context-dependent ones.
发表机构
- Shanghai Jiao Tong University(上海交通大学)
- School of Mathematical Sciences(数学科学学院)
- Institute of Natural Sciences(自然科学研究院)
机构由 AI 辅助整理,请以论文原文为准。