arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

变换秩:架构如何在深度的谱病态中导航

Transforming Rank: How Architecture Navigates the Spectral Pathologies of Depth

Katie Everett

arXiv 2607.14018首次发表:更新:

发表机构

MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究Transformer前馈块架构设计组件在初始化时如何确定跨深度的秩保留,通过重新解释跳跃连接和归一化等,揭示架构各方面对秩的影响,将深度网络架构设计视为在秩崩溃、集成行为和参数数量间的权衡。

AI 中文摘要

我们研究了Transformer前馈块架构设计的每个组件如何在初始化时确定跨深度保留多少秩。我们将长期以来被理解为控制幅度的跳跃连接和归一化重新解释为跨深度保留梯度秩的机制,因为使网络具有表现力的矩阵乘法和非线性激活也会降低秩。我们表明,跳跃连接在秩崩溃和类似集成的行为之间进行权衡,由分支和跳跃的相对比例控制:跳跃连接将梯度绕过秩丢失的残差分支,而不是沿着鼓励层组合的长梯度路径。归一化层的位置通过设置跨深度的分支与跳跃比率来控制相同的权衡,统一了许多归一化位置和深度缩放文献,特别是为什么后归一化会导致秩崩溃而前归一化会使秩平稳。架构的其他方面,如扩展和收缩宽度的双矩阵结构,使用额外参数来保留表示或分支雅可比秩。第二个矩阵去相关一个连贯的平均尖峰,否则它会在具有单个矩阵和非中心激活的块中增长,防止残差表示崩溃。两个矩阵之间的宽度扩展保持分支雅可比满秩:在这个扩展空间中应用降低秩的激活会留下足够的方向来跨越原始空间,宽度遵循马尔琴科 - 帕斯特尔定律。输入 - 输出雅可比的初始化秩预测哪些网络在CIFAR - 10上训练。总之,我们将深度网络的架构设计重新塑造为在秩崩溃、类似集成的行为和参数数量之间进行内在权衡。

英文摘要

We investigate how each component of the Transformer feedforward block architecture design determines how much rank survives across depth at initialization. We reinterpret skip connections and normalization, long understood as controlling magnitude, as mechanisms for preserving gradient rank across depth, since the very matrix multiplications and nonlinear activations that make the network expressive also reduce the rank. We show that skip connections trade off rank collapse against ensemble-like behavior, controlled by the relative scales of the branch and the skip: skip connections route the gradient around the residual branch, where rank is lost, rather than along the long gradient paths that encourage the layers to compose. The placement of the normalization layer controls this same tradeoff by setting the branch-to-skip ratio across depth, unifying much of the normalization placement and depth scaling literature, in particular why rank collapses for Post-Norm but plateaus for Pre-Norm. Other aspects of the architecture, like the two-matrix structure that expands and contracts the width, use additional parameters to preserve the representation or branch Jacobian rank. The second matrix decorrelates a coherent mean spike that would grow across blocks with a single matrix and uncentered activation, preventing the residual representation from collapsing. The width expansion between the two matrices keeps the branch Jacobian full rank: applying the rank-reducing activation in this expanded space leaves enough directions to span the original, at a width that follows a Marchenko--Pastur law. The initialization rank of the input--output Jacobian predicts which networks train on CIFAR-10. Taken together, we recast architecture design for deep networks as navigating an intrinsic tradeoff among rank collapse, ensemble-like behavior, and parameter count.

Comments40 pages. Code: https://github.com/everettk/transforming-rank-paper

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑