发表机构
Seoul National University; SB Intuitions; LG AI Research; Hodoo AI(首尔大学; SB Intuitions; LG AI研究院; Hodoo AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探讨混合Transformer-SSM架构的学习率迁移,发现仅用μP处方即可在多种规模下实现近零LR迁移差距,归因于全局更新-权重不变性与AdamW局部平衡,并验证其跨配置及生产模型的有效性。
AI 中文摘要
我们研究了结合Transformer和状态空间模型(SSM)块的混合架构的学习率(LR)缩放,这类架构已被近期多个生产级语言模型采用。特别地,我们关注在无限宽度、状态大小增长且采用零阶保持(ZOH)离散化条件下为SSM推导的理论缩放规则,与采用简化ZOH Mamba、固定状态大小的领域标准实践实现之间的差距。令人惊讶的是,在这一实践场景中,混合架构在宽度256-2048、深度4-32、直至十亿参数规模下,仅使用原始μP处方即可实现接近零的LR迁移差距,尽管SSM操作超出其张量程序可表示性条件,且我们测试的每种参数化均未通过μP正确性的标准坐标检查诊断。我们将此归因于混合架构中LR迁移的两条件分解:全局更新-权重不变性,由μP的初始化和LR缩放保证;以及局部逐组件平衡,由AdamW的逐参数归一化提供。我们的观察表明,最优LR在宽度上最多8倍内不变,且该宽度不变性在深度、序列长度、批大小和Transformer与SSM比例上均成立,并可迁移至Nemotron-H,一个我们自定义架构集之外的生产级混合模型。我们希望这些发现填补理论缩放规则与实践混合实现之间的差距,并激发进一步研究以弥合这一鸿沟。
英文摘要
We study learning rate (LR) scaling for hybrid architectures combining Transformer and State-Space Model (SSM) blocks, a class adopted by several recent production language models. In particular, we focus on the gap between the theoretical scaling rules derived for SSMs under zero-order-hold (ZOH) discretization at infinite width with growing state size, and the field-standard practical implementations using simplified-ZOH Mamba at fixed state size. Surprisingly, in this practical regime hybrid architectures achieve a near-zero LR transfer gap across widths 256-2048 and depths 4-32 up to billion-parameter scale using only the original $μ$P prescription, even though SSM operations fall outside its Tensor Programs representability conditions and every parameterization we test fails the standard coordinate-check diagnostic of $μ$P correctness. We attribute this to a two-condition decomposition of LR transfer in hybrid architectures: a global update-to-weight invariance, enforced by $μ$P's initialization and LR scaling; and a local per-component balance, provided by AdamW's per-parameter normalization. Our observations show that the optimal LR is invariant to width up to 8$\times$, that this width invariance holds across depth, sequence length, batch size, and Transformer-to-SSM ratio, and that it transfers to Nemotron-H, a production hybrid outside our custom architecture set. We hope these findings fill the gap between theoretical scaling rules and practical hybrid implementations, and stimulate further research toward bridging it.
CommentsAccepted at NeurIPS 2026. 42 pages, 14 figures