稠密混合专家作为重参数化的宽前馈网络:固定计算下的粒度扫描
Dense Mixture-of-Experts as a Reparameterized Wide FFN: A Granularity Sweep at Fixed Compute
浏览论文内容
中文总结 AI 辅助
本文通过固定计算下的稠密MoE粒度扫描,发现K=2的软门控FFN性能最佳,验证损失改善0.0048,并揭示集中路由的功能重要性。
中文摘要 AI 辅助
稀疏混合专家(MoE)模型将学习到的路由与从大型专家池中的选择相结合。我们通过一个稠密类比来隔离动态专家组合的贡献:$K$ 个SwiGLU专家,每个token都激活所有专家,并通过softmax门控组合,在固定的总前馈网络(FFN)宽度下进行。由于没有更大的专家池或离散选择,稠密基线是$K=1$的情况。验证损失随$K$的变化是非单调的:$K=2$比基线改善了$0.0048$,而$K=4$和$K=6$分别使其恶化了$0.0053$和$0.0197$。路由通常保持软性,专家使用均衡,但在$K=4$模型的第一层中,路由几乎是one-hot的。强制该门控均匀化会使诊断子集上的损失增加$2.4$ nats,表明其集中的路由在功能上很重要。我们进一步证明,该架构本质上是一个稠密SwiGLU,具有token依赖的、单纯形约束的组缩放,并且其函数类包含稠密基线。这些单次运行的结果表明,在固定宽度下,始终激活、软门控的FFN存在一个粒度最佳点,其中$K=2$在测试配置中表现最佳。
英文摘要
Sparse Mixture-of-Experts (MoE) models combine learned routing with selection from a large expert pool. We isolate the contribution of dynamic expert combination using a dense analogue: $K$ SwiGLU experts, all active for every token and combined by a softmax gate, at fixed total FFN width. With no larger pool or discrete selection, the dense baseline is the $K=1$ case. Validation loss varies non-monotonically with $K$: $K=2$ improves over the baseline by $0.0048$, whereas $K=4$ and $K=6$ worsen it by $0.0053$ and $0.0197$, respectively. Routing generally remains soft and expert usage balanced, except in the first layer of the $K=4$ model, where routing is nearly one-hot. Forcing this gate to uniform increases loss by $2.4$ nats on a diagnostic subset, indicating that its concentrated routing is functionally important. We further show that the architecture is exactly a dense SwiGLU with token-dependent, simplex-constrained group scaling and contains the dense baseline in its function class. These single-run results suggest a granularity sweet spot for always-active, softly gated FFNs at fixed width, with $K=2$ performing best among the configurations tested.
发表机构
- University of Information Technology(信息技术大学)
- Vietnam National University, Ho Chi Minh City(越南国立大学胡志明市分校)
机构由 AI 辅助整理,请以论文原文为准。