发表机构
Qiuzhen College, Tsinghua University; Yau Mathematical Sciences Center, Tsinghua University; East China Normal University; Beijing Institute of Mathematical Sciences and Applications (BIMSA)(清华大学求真书院; 清华大学丘成桐数学科学中心; 华东师范大学; 北京数学科学与应用研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文从统计学视角将MoE视为局部聚合,推导其风险界并解释门控机制,为理解MoE的设计权衡提供了统一统计框架。
AI 中文摘要
混合专家(MoE)架构通过输入依赖的路由机制组合一组专家预测器来提升模型容量,且通常仅为每个输入激活小部分专家。尽管其在现代大规模模型中愈发重要,但现有理论大多聚焦于参数化或正确指定的MoE模型,对其设计选择(尤其是路由、稀疏激活及共享专家)的统计作用仅部分理解。本文将MoE视为一种局部聚合形式,展示该局部性如何重塑近似-估计-计算的权衡;推导了随专家演化的密集与稀疏路由学习的oracle风险界,分离近似、专家学习及路由估计误差,并刻画稀疏Top-K路由如何在控制单输入计算的同时保留局部聚合的优势。此外,通过输入空间几何解释门控,将路由性能与局部专家优势区域关联,且表明DeepSeekMoE等架构采用的共享专家可提取通用预测结构,使被路由专家聚焦于残差局部变异。综上,这些结果为通过输入依赖的专家聚合理解MoE提供了统一统计框架,其中专家专业化与计算权衡受局部预测结构调控。
英文摘要
Mixture-of-experts (MoE) architectures increase model capacity by combining a collection of expert predictors through input-dependent routing, while often activating only a small subset of experts for each input. Despite their growing importance in modern large-scale models, the statistical roles of their design choices, especially routing, sparse activation, and shared experts, remain only partially understood, as existing theory has largely focused on parametric or correctly specified MoE models. In this paper, we view MoE as a form of localized aggregation and show how this localization reshapes the approximation-estimation-computation tradeoff. We derive oracle risk bounds for learning dense and sparse routing with evolving experts, separating approximation, expert-learning, and router-estimation errors, and characterize how sparse Top-K routing can retain the benefits of localized aggregation while controlling per-input computation. We also interpret gating through the geometry of input space, relating routing performance to regions of local expert advantage, and show how shared experts, as adopted in architectures such as DeepSeekMoE, can extract common predictive structure so that routed experts focus on residual local variation. Together, these results provide a unified statistical framework for understanding MoE through input-dependent expert aggregation, in which expert specialization and computational tradeoffs are governed by local predictive structure.
Comments166 pages, 2 figures