arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

谱异常值揭示Transformer注意力中主导的学习结构

Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention

Kasun Dewage, Marianna Pensky, Suranadi De Silva, T. H. Bandara

arXiv 2608.07921首次发表:更新:

AI 中文总结

本研究运用MP随机矩阵理论分析预训练Transformer注意力权重,发现谱异常值编码学习结构主导成分,通过实验验证其对模型性能的关键作用,相关观察可指导参数高效微调与结构化剪枝。

AI 中文摘要

我们将马尔琴科-帕斯图尔(Marchenko-Pastur, MP)随机矩阵理论应用于预训练的注意力权重,以将每个投影矩阵分离为类随机的主体部分和一组谱异常值。我们通过因果验证了该分解:在Mistral-7B中,将MP识别出的异常值(信号)置零,会使HellaSwag、MMLU和PIQA的性能接近随机水平;而将数量匹配的主体奇异值子集置零,则会导致更小但不可忽略的性能下降。在11个预训练Transformer中,我们识别出5种重复出现的模式:谱异常值编码了学习结构的主导成分;Q投影携带最多的异常值;分组查询注意力中的V投影缺乏清晰的信号/噪声分离;Q中的条目级异常值形成结构化行带,O中形成结构化列带;K和O中特定的残差流维度会跨层持续作为带异常值。最后,我们概述了这些观察结果如何为参数高效微调与结构化剪枝提供指导。

英文摘要

We apply Marchenko-Pastur (MP) random matrix theory to pre-trained attention weights in order to separate each projection matrix into a random-like bulk and a set of spectral outliers. We validate this decomposition causally: zeroing the MP-identified outliers (signal) in Mistral-7B drives HellaSwag, MMLU, and PIQA close to random-chance performance, whereas zeroing a count-matched subset of bulk singular values causes smaller but non-negligible degradation. Across 11 pre-trained transformers we identify five recurring patterns: spectral outliers encode a dominant component of the learned structure; Q projections carry the most outliers; V projections under grouped-query attention lack a clean signal/noise separation; entry-level outliers form structured row-bands in Q and column-bands in O; and specific residual-stream dimensions persist as band outliers across layers in K and O. We close by outlining how these observations could inform parameter-efficient fine-tuning and structured pruning.

CommentsAccepted at the International Conference on Machine Learning and Applications (ICMLA 2026); to appear in IEEE proceedings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑