发表机构
Université de Montréal; Mila – Quebec AI Institute(蒙特利尔大学; 米拉-魁北克人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究运用威尔逊重整化群理论,将Transformer注意力机制视为对MLP残差堆栈不动点的扰动,推导公式并得出预测。通过在合成马尔可夫链序列上测试,发现注意力相关性与数据谱结构有关,且一阶RG扰动框架可解释这种差异。
AI 中文摘要
我们运用威尔逊重整化群理论(RG)的语言,将Transformer的注意力机制视为对训练好的MLP残差堆栈不动点的扰动,探究其是相关、边缘还是无关算子。我们推导了不动点位移公式,并得出关于不动点几何、有效秩分布、层特异性和扰动衰减谱的四个可测试预测。在具有可控相关长度的合成马尔可夫链序列上进行测试,发现:对于长链(长相关),注意力是强相关的,它弥合了MLP无法弥合的残差损失差距并驱动表示空间中的相变;对于短链(短相关),注意力是无关的,Transformer收敛到与MLP相同的损失和不动点几何;过渡由第一层头主导;扰动衰减实验揭示了一种 regime 反转。这些结果表明注意力的相关性不是架构的属性,而是数据生成过程的谱结构的属性,并且一阶RG扰动框架提供了对这种差异的预测性解释。
英文摘要
Using the language of Wilsonian renormalization group theory (RG), we treat the Transformer's attention mechanism as a perturbation of the trained MLP residual-stack fixed point and ask whether it constitutes a relevant, marginal, or irrelevant operator. We derive a fixed-point shift formula and obtain four testable predictions for the fixed-point geometry, effective rank profile, layer specificity, and perturbation decay spectrum. Testing these on synthetic Markov chain sequences with controlled correlation length, we find: (1) For large chains(long correlation), attention is strongly relevant: it closes a residual loss gap the MLP cannot bridge and drives a phase transition in representation space, with effective rank jumping above input dimensionality at layer 1 and stabilizing at a high-dimensional plateau. (2) For short chains(short correlation), attention is irrelevant: the Transformer converges to the same loss and fixed-point geometry as the MLP, though it contracts perturbations faster. (3) The transition is dominated by the first-layer head (L0H0), which accounts for more than 4 times the representational shift of any subsequent head, consistent with the prediction that the relevant operator acts before the MLP begins integrating out positional variation. (4) Perturbation decay experiments reveal a regime reversal: in the long correlation regime the Transformer selectively preserves slow Markov modes (5.4 times the dynamic range in decay length vs. 1.3 times for the MLP); in the short correlation regime it suppresses all modes faster than the MLP, with no spectral selectivity. Together, these results show that the relevance of attention is not a property of the architecture but of the spectral structure of the data-generating process, and that a first-order RG perturbation framework provides a predictive account of that difference.