arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32534cs.CVcs.AIcs.LG

DepthBench:度量残差连接如何实现更深层的有效计算

DepthBench: Measuring How Residual Connections Enable More Computational Depth

Keyu Wang, Yangyi Huang, Jiale Kang, David González-Martínez, Weiyang Liu, Shiwei Liu

AI总结:

本文提出DepthBench基准,通过控制模型规模和预训练设置、变化宽深比,在10种架构上系统评估残差连接对计算深度的影响,发现残差连接设计是决定深度能否有效扩展的关键因素。

AI中文摘要:

深度是增加Transformer计算容量的自然方式,然而随着深度增大,更深层的贡献可能会减弱。近期的方法通过增强归一化(如LayerNorm Scaling)或残差连接(如mHC、AttnRes)来改善信息流动和深度利用。然而,目前尚不清楚这些方法是否真正将增加的架构深度转化为有效的计算深度,以及其报告的性能提升究竟源于跨深度更好的信息访问,还是未计入的混淆因素。本文提出了DepthBench,一个用于研究多种架构下计算深度的受控基准。我们系统地改变宽度-深度纵横比(d_model/n_layer),从浅宽形状到深窄形状,同时保持模型规模和预训练设置固定。在10种代表性架构上,我们发现将更多容量分配给深度的收益高度依赖于架构。标准的Pre-LN及其大多数基于归一化和缩放的变体收益甚微,甚至随着模型变深变窄而性能下降,而HC和Full AttnRes即使在极端深窄形状下也持续改进。这些收益不仅体现在预训练损失上,还持续转化为改进的领域特定性能和有效计算。受控的逐层分析进一步表明,HC和Full AttnRes的收益与更有效地利用额外层相关,揭示了不同架构间计算深度的不同机制。总体而言,我们的结果识别出残差连接设计是深度能否作为有意义扩展轴的关键决定因素,通过使额外的架构深度转化为有效计算来实现。

英文摘要:

Depth is a natural way to increase the computational capacity in Transformers, yet the contribution of deeper layers can diminish as depth grows larger. Recent approaches enhance normalization (\text{e.g.}, LayerNorm Scaling) or residual connections (\text{e.g.}, mHC, AttnRes) to enable better information flow and depth utilization. However, it remains unclear whether they truly translate increased architectural depth into effective computational depth, and whether their reported gains stem from better access to information across depth, or unaccounted-for confounding factors. In this paper, we introduce \textbf{DepthBench}, a controlled benchmark for studying computational depth across various architectures. We systematically vary the width--depth aspect ratio ($d_{\text{model}}/n_{\text{layer}}$) from shallow--wide to deep--narrow shapes, while keeping the model size and pre-training recipe fixed. Across 10 representative architectures, we find that the benefit of allocating more capacity to depth is strongly architecture-dependent. Standard Pre-LN and most of its norm- and scaling-based variants provide little benefit and can even degrade performance as models become deeper and narrower, whereas HC and Full AttnRes improve consistently even at extreme deep shapes. These gains extend beyond pre-training loss and consistently translate into improved domain-specific performance and effective computation. Controlled layer-level analyses further show that the gains of HC and Full AttnRes are associated with more effective utilization of additional layers, revealing distinct mechanisms of computational depth across architectures. Overall, our results identify residual connection design as a key determinant of whether depth can serve as a meaningful scaling axis by enabling additional architectural depth to translate into effective computation.

↑