发表机构
College of Computer Science, Beijing University of Technology; School of Mathematics, Statistics and Mechanics, Beijing University of Technology(北京工业大学计算机学院; 北京工业大学数学、统计与力学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现Transformer注意力中存在缩放幂等性的代数规律,通过OV几何分解与值共享扩展,揭示了训练注意力头的稀疏方向特性及局部算子代数关系。
AI 中文摘要
我们在Transformer注意力中发现了一种循环代数规律性:有效OV算子T=OV^T的稀疏子集在复合运算下几乎闭合,即T²≈αT。在6个预训练端点(参数规模28亿至2350亿)中,有3.98%至8.00%的注意力头达到平方闭合对齐度P≥0.9,而同一层内未发现匹配的O/V不匹配情况。精确的主坐标分解T=Q_OKQ_V^T和T²=Q_O(KDK)Q_V^T将支持域内传输与读写返回几何分离。在9个MHA/GQA模型的7304个注意力头中,仅打乱K的方向(同时保留奇异值、范数、因子跨度和主角度)会使中位数闭合度从0.336降至1.04×10^-4;训练得到的方向在98.64%的注意力头及所有层中均表现更优。构造性搜索表明,所有被考察层均具备实现高闭合度的潜力,但通常未被达到。在三个独立训练谱系中的回溯轨迹进一步区分了广泛可用的容量与最终强注意力头所获得的方向。在精确值共享下,注意力头级闭合度扩展为右作用代数T_iT_j=α_jT_i。7个模型的实验验证了该近似定律,并揭示了具有共享值定义核的不同斜投影。这些结果将缩放幂等性表征为广泛可用几何容量内的稀疏训练方向,并展示了值共享如何将注意力头级关系扩展为局部算子代数。
英文摘要
We identify a recurrent algebraic regularity in Transformer attention: a sparse subset of effective OV operators $T=OV^\top$ nearly closes under composition, $T^2\approxαT$. Across six pretrained endpoints spanning 2.8B--235B parameters, 3.98--8.00% of heads reach squared closure alignment $\mathcal{P}\geq0.9$, while no matched within-layer O/V mismatch does. An exact principal-coordinate factorization, $T=Q_OKQ_V^\top$ and $T^2=Q_O(KDK)Q_V^\top$, separates within-support transport from read--write return geometry. Across all 7,304 heads in nine MHA/GQA models, scrambling only the orientation of $K$ while preserving singular values, norms, factor spans, and principal angles reduces median closure from 0.336 to $1.04\times10^{-4}$; trained orientation wins for 98.64% of heads and in every layer. Constructive searches show that high closure is feasible in every surveyed layer, but usually not attained. Retrospective trajectories in three independently trained lineages further separate broadly available capacity from the orientations attained by final strong heads. Under exact value sharing, headwise closure extends to a right-action algebra, $T_iT_j=α_jT_i$. Seven-model experiments verify the approximate law and reveal distinct oblique projections with a shared value-defined kernel. These results characterize scaled idempotence as a sparse trained orientation within broadly available geometric capacity and show how value sharing extends a headwise relation into a local operator algebra.
Comments14 pages, 2 figures. Preprint