arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

瞬态储备、汇阻尼器以及注意力传播器中特征值推理的失效

Transient Reserves, Sink Dampers, and the Failure of Eigenvalue Reasoning in the Attention Propagator

Li Hengyu

arXiv 2607.09279首次发表:更新:

AI 中文总结

研究因果变压器注意力矩阵非正规性,通过预解式观点测试其对已训练变压器的预测。利用掩码和投影器分析,发现学习到的非正规性有符号,路由少数群体有瞬态储备,汇多数群体起阻尼作用,特征值预测误差大,预解式特征对相关特性至关重要。

AI 中文摘要

因果变压器的注意力矩阵是行随机的,按深度迭代且本质上是非正规的。对于非正规算子,特征值仅控制渐近行为,有限深度行为由诸如伪谱和克雷斯常数等预解式量控制。我们在预先设定的标准下测试这种预解式观点是否能预测特征值所遗漏的关于已训练变压器的信息。两个结构事实组织了分析:掩码将每个因果随机矩阵的克雷斯常数固定为\(\sqrt{n}\),对掩码强制的佩龙投影器进行收缩可将深度偏差动力学精确分解为收缩算子的乘积。在GPT - 2、Pythia - 410m和Llama - 3 - 8B中,学习到的非正规性被证明是有符号的。一个路由少数群体携带多余的瞬态储备,其跟踪前一个 token 的功能并在归纳头参与时翻倍,而汇多数群体被抑制到匹配洗牌零值以下,因此注意力汇充当瞬态阻尼器。在深度乘积上,存活偏差的特征值预测误差达七到十一个数量级,而在匹配零值中不存在此误差。检查点普查将这种组织追溯到电路形成后的巩固阶段,对Llama - 3 - 8B的钳位干预建立了从三个大规模激活维度通过汇注意力到瞬态阻尼的因果链;LayerNorm模型在其他地方实现相同功能。交叉验证竞赛得出结论,预解式特征对于深度瞬态持久性和路由头识别是必需的,并且任何单一算子的总结都无法预测每个头的因果关键性。

英文摘要

The attention matrix of a causal transformer is row-stochastic, iterated over depth, and non-normal by construction. For non-normal operators, eigenvalues control only asymptotic behavior; finite-depth behavior is controlled by resolvent quantities such as pseudospectra and Kreiss constants. We test, under pre-registered criteria, whether this resolvent view predicts anything about trained transformers that eigenvalues miss. Two structural facts organize the analysis: the mask pins the Kreiss constant of every causal stochastic matrix at $\sqrt{n}$, and deflating the mask-forced Perron projector factorizes the depth deviation dynamics exactly into a product of deflated operators. Across GPT-2, Pythia-410m, and Llama-3-8B, learned non-normality proves to be signed. A routing minority carries excess transient reserve that tracks previous-token function and doubles when induction heads engage, while the sink majority is suppressed below matched shuffle nulls, so that attention sinks act as transient dampers. On depth products, eigenvalue predictions of surviving deviations err by seven to eleven orders of magnitude, an error absent in matched nulls. Checkpoint censuses date this organization to a consolidation phase after circuit formation, and a clamping intervention on Llama-3-8B establishes a causal chain from three massive activation dimensions through sink attention to transient damping; LayerNorm models implement the same functions elsewhere. A cross-validated contest concludes that resolvent features are required for depth-transient persistence and routing-head identity, and that no single-operator summary of any kind predicts per-head causal criticality.

Comments17 pages, 8 figures. Companion paper: arXiv:2607.06621

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑