arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00748cs.CL

测量Transformer深度中的最优传输

Measuring Optimal Transport in Transformer Depth

Alexandre Quemy

首次发表
浏览论文内容

中文总结 AI 辅助

该研究测量Pythia-160m与Pythia-410m在Transformer各层的点云移动是否符合最优传输规律,发现最后一层一致性随训练增强,中间多层块成本接近最优。

中文摘要 AI 辅助

Transformer将每个token的状态从一层传递到另一层,所有词汇共同形成一个随深度变化的点云。我们探究经过训练的网络是否会以最优传输的方式移动该点云:以最低成本,沿将每个token与其最优目标配对的映射移动。我们在Pythia-160m和Pythia-410m上进行测量,采用连续层点云之间的精确分配、测量得到的采样下限、对已知最优耦合的校准,以及将成本拆分为点云的共同偏移和token特定移动的方法。在最后一层,两个模型都将token移动到最优传输映射指定的位置,其中Pythia-410m达到最优成本,Pythia-160m则略高于最优成本;在第一层,它们未做到这一点。在中间层,十次转换中仅有两次可按成本判断单个层,而多层块移动点云的成本接近最优。初始化时最后一层的一致性远弱于训练后(0.64对比0.86),且随训练增强。

英文摘要

A transformer carries each token's state from layer to layer, and the whole vocabulary carried together forms a cloud that moves with depth. We ask whether a trained network moves this cloud the way optimal transport would: at the cheapest cost, and along the map that pairs each token with its optimal destination. We measure both on Pythia-160m and Pythia-410m, with an exact assignment between consecutive layer clouds, a measured sampling floor, calibration on couplings known to be optimal, and a split of the cost into the common shift of the cloud and the token-specific moves. At the last layer, both models move their tokens where the optimal-transport map sends them, at the optimal cost for Pythia-410m and slightly above it for Pythia-160m. At the first layer they do not. In between, single layers can be judged on cost at only two of ten transitions, and blocks of several layers move the cloud at close to the optimal cost. The agreement at the last layer is much weaker at initialisation (0.64 against 0.86) and grows with training.

发表机构

  • Hother Labs(霍瑟实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑