arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MASKerade:用于稠密到MoE升级的令牌路由掩码专家

MASKerade: Token-Routed Mask Experts for Dense-to-MoE Upcycling

Mingyuan Zhang, Yue Bai, Zhongruo Wang, Yupin Huang, Yiyang Huang, Hailing Wang, Huimin Zeng, Yun Fu

arXiv 2610.07809首次发表:更新:

发表机构

Northeastern University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MASKerade通过冻结FFN上的学习掩码定义专家,实现稠密到MoE升级,在视觉语言基准上以较低计算成本达到最优性能。

AI 中文摘要

稀疏激活的混合专家(MoE)模型在不按比例增加每令牌计算量的情况下增加了模型容量。稠密到MoE升级通过重用预训练的稠密模型来构建此类系统,通常是将前馈网络(FFN)复制到独立训练的专家中。我们引入了MASKerade,一种稠密到MoE的训练方法,该方法将专家学习为冻结的预训练FFN的稀疏子网络。每个专家由一个学习得到的二值掩码定义,一个令牌级路由器选择要执行和组合的掩码FFN。路由器和掩码分数联合优化,而底层FFN权重值保持不变。这种公式支持在同一路由架构内的神经元结构化、半结构化和非结构化专家。我们的主要配置使用四个2:4专家和top-2路由,其中两次半稠密专家传递具有名义上的一次稠密传递的FFN算术,无需独立的专家权重矩阵。在Qwen和Gemma骨干的五个视觉语言基准上,该配置在比较的基线中实现了最高性能。跨掩码粒度、路由干预和计算匹配对照的比较区分了学习连接与专家激活次数的影响。这些结果确立了在冻结权重上进行掩码学习作为构建令牌路由MoE专家的实用替代方案。

英文摘要

Sparsely activated Mixture-of-Experts (MoE) models increase model capacity without a proportional increase in per-token computation. Dense-to-MoE upcycling reuses pretrained dense models to construct such systems, commonly by copying feed-forward networks (FFNs) into independently trained experts. We introduce MASKerade, a dense-to-MoE training method that instead learns experts as sparse subnetworks of a frozen pretrained FFN. Each expert is defined by a learned binary mask, and a token-level router selects which masked FFNs to execute and combine. The router and mask scores are optimized jointly, while the underlying FFN weight values remain unchanged. This formulation supports neuron-structured, semi-structured, and unstructured experts within the same routing architecture. Our main configuration uses four 2:4 experts with top-2 routing, where two half-dense expert passes have the nominal FFN arithmetic of one dense pass, without requiring independent expert weight matrices. On five vision-language benchmarks with Qwen and Gemma backbones, this configuration achieves the highest performance among the compared baselines. Comparisons across mask granularities, routing interventions, and compute-matched controls distinguish the effects of learned connectivity from expert activation count. These results establish mask learning over frozen weights as a practical alternative for constructing token-routed MoE experts.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑