发表机构
University of Luxembourg; London Institute for Mathematical Sciences(卢森堡大学; 伦敦数学科学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过比较中间与最终残差状态,揭示Transformer推理中表示几何特异性的形成机制,并证明直线收敛无法解释观察到的竞争变化。
AI 中文摘要
Transformer语言模型通过连续的残差更新来构建预测,但其表示如何变得特定于最终结果仍不清楚。我们通过将中间残差状态与其自身的最终状态以及来自其他上下文的经验性最终状态库进行比较来研究这一过程。在六个预训练语言模型中,自身终点相对于平均替代终点更早变得可取,而许多单个终点仍然更接近。这些竞争集合通常随深度增加而缩小,但其成员会发生变化,且幸存的终点之间不必变得更加相似。因此,方向对齐和终点排名可以改善,而到最终状态的欧几里得距离变化很小。我们开发了一个简单的高维模型,分离了范数、对齐和终点几何的作用,展示了逐渐的方向变化如何产生竞争性的急剧减少。我们还证明了,在欧几里得距离或余弦距离下,朝向自身终点的直线路径不能引入新的竞争者;因此,观察到的条目确立了与直线收敛的偏离。最后,与排名较低的输出标记相关联的终点在所有研究的模型中往往在余弦距离上更远,将残差几何与输出组织联系起来。总之,这些发现刻画了Transformer推理过程中几何特异性的增加,并解释了为什么距离、竞争者数量以及幸存终点的集中度提供了对该过程的不同视角。
英文摘要
Transformer language models build predictions through successive residual updates, but how their representations become specific to an eventual outcome remains unclear. We study this process by comparing intermediate residual states with their own final states and an empirical bank of final states from other contexts. Across six pretrained language models, the own endpoint becomes preferable to the average alternative early, while many individual endpoints remain closer. These competing sets generally shrink with depth, but their membership changes and their surviving endpoints need not become more similar to one another. Directional alignment and endpoint rank can therefore improve while Euclidean distance to the final state changes little. We develop a simple high-dimensional model that separates the roles of norm, alignment, and endpoint geometry, showing how gradual directional changes can produce sharp reductions in competition. We also prove that a straight path toward the own endpoint cannot introduce new competitors under either Euclidean or cosine distance; observed entries thus establish departures from straight-line convergence. Finally, endpoints associated with lower-ranked output tokens tend to lie farther away in cosine distance across all studied models, connecting residual geometry to output organization. Together, these findings characterize increasing geometric specificity during transformer inference and explain why distance, competitor count, and concentration of the surviving endpoints provide distinct views of that process.