arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

冗余与协同:基于子模优化的MoE依赖感知专家选择

Redundancy Meets Synergy: Dependency-aware Expert Selection for MoE via Submodular Optimization

Zheng Lin, Shaoke Fang, Yuxin Zhang, Jinfeng Xu, Zihan Fang, Zhe Chen, Wei Ni, Jun Luo, Symeon Chatzinotas

arXiv 2610.00558首次发表:更新:

发表机构

Interdisciplinary Centre for Security, Reliability and Trust, University of Luxembourg; Department of Computer Science, Peking University; Institute of Space Internet, Fudan University; Department of Electrical and Computer Engineering, The University of British Columbia; Department of Computer Science, City University of Hong Kong; School of Engineering, Edith Cowan University; College of Computing and Data Science, Nanyang Technological University(卢森堡大学安全、可靠性与信任跨学科中心; 北京大学计算机科学系; 复旦大学空间互联网研究院; 不列颠哥伦比亚大学电气与计算机工程系; 香港城市大学计算机科学系; 埃迪斯科文大学工程学院; 南洋理工大学计算与数据科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对MoE专家选择忽视依赖关系的问题,提出DS-MoE框架,利用差分子模优化解耦冗余与协同,并设计MM算法高效选择专家子集,实验验证其性能优于现有基线。

AI 中文摘要

尽管混合专家(Mixture-of-Experts, MoE)模型通过稀疏激活有效扩展了模型容量,但其部署常因高昂的内存需求而受阻。提取一个紧凑的专家子集是一种有前景的解决方案。然而,现有的专家选择启发式方法主要依赖Top-k排序,这种排序孤立地评估单个专家,忽略了MoE门控网络引入的复杂专家间依赖关系。在本文中,我们提出DS-MoE,一个基于理论框架的方法,通过差分子模(difference-of-submodular, DS)优化重新定义专家选择。通过分析损失衰减的二阶泰勒展开,我们揭示了专家组合中的功能二元性:冗余(专家编码重叠表示)和协同(专家提供互补的错误抵消)。为应对这种二元性,我们通过将选择目标表述为DS函数,从数学上将冗余减少与协同最大化解耦。此外,我们设计了一种定制的最大化-最小化(majorization-minimization, MM)算法,具有可证明的单调性保证,以高效识别最优专家子集。大量实验表明,DS-MoE有效保留了不可或缺的专家组合,与最先进的基线相比取得了更优的性能。

英文摘要

While Mixture-of-Experts (MoE) models effectively scale model capacity through sparse activation, their deployment is often bottlenecked by prohibitive memory requirements. Extracting a compact subset of experts presents a promising solution. However, existing expert selection heuristics predominantly rely on Top-k ranking, which isolates the evaluation of individual experts and ignores the intricate inter-expert dependencies introduced by the MoE gating network. In this paper, we propose DS-MoE, a theoretically grounded framework that redefines expert selection via difference-of-submodular (DS) optimization. By analyzing the second-order Taylor expansion of the loss degradation, we reveal functional duality within expert combinations: redundancy (where experts encode overlapping representations) and synergy (where experts provide complementary error cancellation). To navigate this duality, we mathematically decouple redundancy reduction from synergy maximization by formulating the selection objective as a DS function. Furthermore, we devise a tailored majorization-minimization (MM) algorithm with provable monotonicity guarantees to efficiently identify the optimal expert subset. Extensive experiments demonstrate that DS-MoE effectively preserves indispensable expert combinations, achieving superior performance compared to the state-of-the-art baselines.

Comments26 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑