arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SMELT:计算匹配型MoE循环Transformer的缩放定律

SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

Shaowen Wang, Ge Zhang, Kairong Luo, Yuhao Wu, Shaofan Liu, Jiaheng Liu, Wenhao Huang, Shen Yan, Jian Li

arXiv 2609.01343首次发表:更新:

发表机构

Tsinghua University; ByteDance Seed; M-A-P(清华大学; 字节跳动Seed; M-A-P)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究在严格匹配关键预算指标的前提下,提出SMELT方案,通过循环MoE Transformer中间层,提升了模型性能并节省了训练FLOPs,为Transformer架构优化提供了实用方案。

AI 中文摘要

循环Transformer通过迭代共享层块来增加有效深度,但多数评估在固定模型规模下进行,将架构优势与额外浮点运算量(FLOPs)混淆。我们在严格匹配每token的FLOPs、总非嵌入参数数量及键值(KV)缓存的前提下,研究混合专家(Mixture-of-Experts, MoE)Transformer的循环设计。通过一系列消融实验,我们提出了名为SMELT(Sparse MoE Transformer,中间层循环两次)的方案:将中间一半的层循环两次,同时在上述三个预算指标上与未循环的基准模型(Baseline)匹配。我们将SMELT扩展至四个规模,最大非嵌入参数达540亿,并为每种架构拟合了独立的Chinchilla式缩放定律。SMELT的损失随计算量下降更快,在计算最优前沿可节省6.8%至18.0%的训练FLOPs。该优势可迁移至下游基准任务,且超出验证损失的预测效果,在代码任务上表现最优,并随样本长度和上下文示例数量增加而提升。机制分析表明,第二次循环可减少注意力汇聚(attention sink),将注意力权重更多导向与内容相关的token,这种归纳偏置可能是观测到性能提升的原因。这些结果表明,即使在预算匹配的条件下,循环设计也能改进Transformer,提供了一种将深度复用转化为可衡量收益的实用方案。

英文摘要

Looped Transformers increase effective depth by iterating a shared block of layers, but most evaluations compare at fixed model size, conflating architectural advantage with extra FLOPs. We study looping on Mixture-of-Experts Transformers while closely matching per-token FLOPs, total non-embedding parameters, and KV cache. Through a series of ablations, we arrive at a recipe we call SMELT (Sparse MoE Transformer, middle layers Loop Twice), which loops the middle half of layers twice while matching the unlooped Baseline on all three budgets. We scale SMELT across four sizes up to 54B non-embedding parameters and fit a separate Chinchilla-style scaling law for each architecture. SMELT's loss drops faster with compute, saving 6.8--18.0\% of training FLOPs on the compute-optimal frontier. The advantage transfers to downstream benchmarks beyond what validation loss predicts, is largest on Code, and grows with sample length and the number of in-context examples. Mechanistic analysis shows that the second visit reduces the attention sink and redirects mass toward content-relevant tokens, an inductive bias that may underlie the observed performance gains. These results show that looping can improve Transformers even under budget matching, offering a practical recipe that turns depth reuse into measurable gains.

Comments36 pages, 25 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑