发表机构
Hother Labs(霍瑟实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究以Pythia-160M和Pythia-410M为对象,用离散朗之万模型表征Token表示的深度流,发现其非线性、不下降自身密度且含不可忽略的旋转分量,揭示了Token不同排名在网络层的变化规律。
AI 中文摘要
Token的表示会逐层通过网络,所有Token共同构成一个流。我们针对Pythia-160M和Pythia-410M的语料库平均轨迹,将该流的运动方程拟合为离散朗之万模型,并在保留的Token上对预测步骤进行评分。线性映射常被用作层的廉价替代,它们所概括的流并非线性:在两个模型的每一次转换中,二次漂移都优于线性映射,且当Kramers-Moyal估计器的邻域保持局部时,其结果一致。我们进一步对该流进行表征:首先,证明它不会下降其自身的对数密度,漂移反而下降的是一个非密度的势;其次,旋转分量不可忽略,占可解释漂移的4%至45%,环流体现在流所保留的内容:Token在全部13层中保持其角排名,而其范数排名被打乱,浓度排名则在最后一个模块中反转。
英文摘要
A token's representation is carried through the network layer by layer. The whole vocabulary carried together forms a flow. We fit this flow's equation of motion as a discrete Langevin model over corpus-mean trajectories of Pythia-160M and Pythia-410M, and score the predicted steps on held-out tokens. Linear maps are often used as cheap surrogates for a layer. The flow they summarize is not linear: a quadratic drift beats the linear linear map at every transition of both models, and the Kramers--Moyal estimator agrees wherever its neighborhoods stay local. We then characterize the flow further. First, we show that it does not descend its own log-density. The drift instead descends a potential that is not the density. Second, the rotational component is not negligible, $4$ to $45\%$ of the explainable drift, and the circulation shows in what the flow preserves: a token keeps its angular rank across all thirteen layers while its norm rank is shuffled and its concentration rank is reversed by the last block.