arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

X-MoD:超越混合深度(Mixture-of-Depths)的稀疏深度路由实用缩放定律

X-MoD: Practical Scaling Laws for Sparse-Depth Routing Beyond Mixture-of-Depths

Bowen Dong, Yilong Fan, Tengyu Pan, Yike Zhang, Zhenyu Li, Zijian Zhang, Xuewei Li, Mei Yu, Jianyong Wang

arXiv 2609.34212首次发表:更新:

发表机构

Tsinghua University; Tianjin University(清华大学; 天津大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

X-MoD提出一种解耦令牌稀疏性与锚点步长的稀疏深度架构,并建立实用缩放定律,以预测和优化不同配置下的性能,超越传统MoD。

AI 中文摘要

混合深度(MoD)通过仅将一部分令牌路由到选定的层,实现了跨Transformer深度的条件计算,但其原始的“一稀疏一密集”交替模式将总容量与活跃容量紧密耦合,限制了稀疏深度的扩展。我们提出了X-MoD,一种可扩展的稀疏深度架构,它将令牌稀疏性与锚点步长解耦,允许总参数量增长,同时保持活跃等效容量几乎不变。为了使深层稀疏路由可训练,X-MoD结合了密集锚点、方差缩放的逐层门控和深度方向的令牌平衡。为了使这一机制可分析且可用,我们将稀疏深度路由形式化为一个条件架构设计问题:给定计算量、上下文长度和活跃等效骨干规模,应如何选择路由配置?我们通过将X-MoD相对于FLOP匹配的密集基线进行拟合,开发了一个实用的缩放定律框架,产生了一个可解释的定律,将性能分解为稀疏容量增益、稀疏上下文校正和锚点步长交互。该定律可预测各种路由配置下的验证损失,并揭示上下文长度、模型规模和锚点步长如何影响稀疏深度性能。我们通过预训练扫描、留出缩放定律预测、消融实验、下游评估以及与Dense、MoD和代表性MoE基线的比较,验证了该架构和定律。

英文摘要

Mixture-of-Depths (MoD) enables conditional computation across Transformer depth by routing only a subset of tokens through selected layers, but its original one-sparse--one-dense alternation tightly couples total capacity to active capacity and limits sparse-depth scaling. We introduce X-MoD, a scalable sparse-depth architecture that decouples token sparsity from anchor stride, allowing total parameter count to grow while keeping active-equivalent capacity nearly fixed. To make deep sparse routing trainable, X-MoD combines dense anchors with variance-scaled layer-wise gating and depth-wise token balancing. To make this regime analyzable and usable, we formulate sparse-depth routing as a conditional architecture-design problem: given compute, context length, and active-equivalent backbone size, how should the routing configuration be chosen? We develop a practical scaling-law framework by fitting X-MoD relative to FLOP-matched dense baselines, yielding an interpretable law that decomposes performance into sparse-capacity gain, sparse-context correction, and anchor-stride interaction. The law predicts validation loss across routing configurations and reveals how context length, model scale, and anchor stride shape sparse-depth performance. We validate the architecture and law through pretraining sweeps, held-out scaling-law prediction, ablations, downstream evaluations, and comparisons with Dense, MoD, and representative MoE baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑