学习率迁移用于混合Transformer-SSM架构
Learning Rate Transfer for Hybrid Transformer-SSM Architectures
浏览论文内容
中文总结 AI 辅助
本研究探讨混合Transformer-SSM架构的学习率迁移,发现仅用μP处方即可在多种规模下实现近零LR迁移差距,归因于全局更新-权重不变性与AdamW局部平衡,并验证其跨配置及生产模型的有效性。
中文摘要 AI 辅助
我们研究了结合Transformer和状态空间模型(SSM)块的混合架构的学习率(LR)缩放,这类架构已被近期多个生产级语言模型采用。特别地,我们关注在无限宽度、状态大小增长且采用零阶保持(ZOH)离散化条件下为SSM推导的理论缩放规则,与采用简化ZOH Mamba、固定状态大小的领域标准实践实现之间的差距。令人惊讶的是,在这一实践场景中,混合架构在宽度256-2048、深度4-32、直至十亿参数规模下,仅使用原始μP处方即可实现接近零的LR迁移差距,尽管SSM操作超出其张量程序可表示性条件,且我们测试的每种参数化均未通过μP正确性的标准坐标检查诊断。我们将此归因于混合架构中LR迁移的两条件分解:全局更新-权重不变性,由μP的初始化和LR缩放保证;以及局部逐组件平衡,由AdamW的逐参数归一化提供。我们的观察表明,最优LR在宽度上最多8倍内不变,且该宽度不变性在深度、序列长度、批大小和Transformer与SSM比例上均成立,并可迁移至Nemotron-H,一个我们自定义架构集之外的生产级混合模型。我们希望这些发现填补理论缩放规则与实践混合实现之间的差距,并激发进一步研究以弥合这一鸿沟。
英文摘要
We study learning rate (LR) scaling for hybrid architectures combining Transformer and State-Space Model (SSM) blocks, a class adopted by several recent production language models. In particular, we focus on the gap between the theoretical scaling rules derived for SSMs under zero-order-hold (ZOH) discretization at infinite width with growing state size, and the field-standard practical implementations using simplified-ZOH Mamba at fixed state size. Surprisingly, in this practical regime hybrid architectures achieve a near-zero LR transfer gap across widths 256-2048 and depths 4-32 up to billion-parameter scale using only the original $μ$P prescription, even though SSM operations fall outside its Tensor Programs representability conditions and every parameterization we test fails the standard coordinate-check diagnostic of $μ$P correctness. We attribute this to a two-condition decomposition of LR transfer in hybrid architectures: a global update-to-weight invariance, enforced by $μ$P's initialization and LR scaling; and a local per-component balance, provided by AdamW's per-parameter normalization. Our observations show that the optimal LR is invariant to width up to 8$\times$, that this width invariance holds across depth, sequence length, batch size, and Transformer-to-SSM ratio, and that it transfers to Nemotron-H, a production hybrid outside our custom architecture set. We hope these findings fill the gap between theoretical scaling rules and practical hybrid implementations, and stimulate further research toward bridging it.
发表机构
- Seoul National University(首尔大学)
- SB Intuitions
- LG AI Research(LG AI研究院)
- Hodoo AI
机构由 AI 辅助整理,请以论文原文为准。