AI 中文总结
本文统一十四篇研究,发现SVD压缩注意力投影对秩坍缩呈相位反转:初始化时抑制、预训练时加速,归因于子空间选择而非范数缩减。
AI 中文摘要
线性代数提供了现代人工智能通过神经网络编码、压缩和传播信息所采用的概念框架(矩阵秩、奇异值分解(SVD)和特征分解)。本文统一了十四篇独立的同行评审工作,分析这些技术在基于Transformer的基础模型研究中的使用情况,聚焦于该主题的三个领域:自注意力矩阵输出秩的推导与性质、有意利用该现象的压缩方法,以及低秩键值(KV)缓存投影及其与线性注意力和状态空间结构化模型的半可分矩阵对偶性。我们是在观察到该文献中的一个开放问题后开展本工作的:上述压缩方法与网络自然秩坍缩之间的相互作用。在本文中,我们报告了一个原始发现:对注意力投影使用SVD压缩实际上对网络秩坍缩具有相反的效果——虽然它在初始化时强烈抑制秩坍缩,但在预训练模型(GPT-2 124M、GPT-2 Medium 355M和Pythia-160M)上却加速了秩坍缩,且对象混叠伪影出现的风险极小(在所有压缩比下均验证),并在四种秩估计方法中保持一致。对两种设置中该效应的受控因果分解表明,这种行为的原因可以解释为SVD在压缩矩阵时对子空间的选择优于其实现的算子范数缩减,这解释了初始化时约76%的效应和预训练权重上约83%的效应,为校准感知压缩观点提供了细化,并解释了其为何优于朴素SVD截断。
英文摘要
Linear algebra provides the framework of concepts (matrix rank, singular value decomposition (SVD), and eigendecomposition) that modern artificial intelligence employs to encode, compress, and propagate information through neural networks. This paper unifies fourteen separate peer-reviewed works analyzing the usage of these techniques in the context of transformer-based foundation model research, focusing on three areas of the topic: derivations and properties of self-attention matrices' output rank, compression methods that purposefully utilize this phenomenon, and the low-rank key-value (KV) cache projection and its semiseparable-matrix duality to linear attention and state-space structured models. We were motivated to conduct this work after observing an open problem in this literature: the interplay of the mentioned compression methods with natural rank collapse of the network. With this paper, we report an original finding that using SVD compression of attention projections actually has the opposite effect on the rank collapse of the network: while it strongly suppresses it at initialization, it accelerates on pretrained models (for GPT-2 124M, GPT-2 Medium 355M, and Pythia-160M) with minimal risk of object aliasing artifacts appearing (verified on all compression ratios) and is consistent across four rank estimation methods. A controlled causal decomposition of the effect in both settings showed that the reason for this behavior can be explained by the choice of the subspace SVD makes when compressing the matrix better than the reduction of the operator norm it achieves, explaining roughly 76% of the effect at initialization and 83% on the pretrained weights, providing a refinement to the calibration-aware compression viewpoint and an explanation of why it outperformed naive SVD truncation.