元学习专家分配位置:面向混合专家模型的任务条件化分层压缩
Meta-Learning Where to Allocate Experts: Task-Conditioned Layer-Wise Compression for MoEs
浏览论文内容
中文总结 AI 辅助
该研究提出MetaNet控制器,为MoE模型分层预测专家分配阈值,在DeepSeek-MoE-16B-Chat上实现准确率与专家激活的权衡,且控制器可跨数据集迁移。
中文摘要 AI 辅助
混合专家(MoE)模型将每个token路由到一部分专家网络,在保持每个token计算稀疏性的同时提升模型容量。在许多已部署的MoE模型中,各层和任务的活跃专家数量是固定的,但层的作用随深度变化,专家冗余度也随深度变化,而活跃专家的需求随任务难度变化。现有方法仅解决该场景的部分问题:分层分配通常离线确定并复用至所有任务,而token级方法使用局部路由信号改变专家激活,未结合任务级上下文。我们提出MetaNet,一种支持集控制器,可为每一层预测专家保留阈值和有界路由偏置,主干网络、专家和路由器保持冻结。在DeepSeek-MoE-16B-Chat上,MetaNet提供可调节的准确率-专家激活权衡:相对于固定k=6的设置,保守设置平均激活3.61个专家(减少40%),MMLU准确率相当(0.489 vs 0.474);激进设置平均激活2.28个专家(减少62%),准确率约低3.7个百分点。经MMLU训练的控制器无需重新训练即可迁移至C-Eval,平均激活2.90个专家(较固定k=6减少52%),准确率为0.386。
英文摘要
Mixture-of-Experts (MoE) models route each token to a subset of expert networks, increasing capacity while keeping per-token computation sparse. In many deployed MoEs, the number of active experts is fixed across layers and tasks, although layer roles and expert redundancy vary with depth and demand varies with difficulty. Existing approaches address only part of this setting: layer-wise allocations are usually determined offline and reused for all tasks, while token-level methods vary expert activation using local routing signals without task-level context. We propose MetaNet, a support-set controller that predicts, for each layer, an expert-retention threshold and a bounded routing bias. The backbone, experts, and router remain frozen. On DeepSeek-MoE-16B-Chat, MetaNet provides a tunable accuracy-expert-activation trade-off. Relative to fixed k=6, a conservative setting activates 3.61 experts on average (40% fewer) and achieves comparable MMLU accuracy (0.489 vs. 0.474), whereas an aggressive setting activates 2.28 experts on average (62% fewer) with accuracy approximately 3.7 percentage points lower. The MMLU-trained controller also transfers to C-Eval without retraining, activating 2.90 experts on average (52% fewer than fixed k=6) at 0.386 accuracy.
发表机构
- Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
- Nanjing College, University of Chinese Academy of Sciences(中国科学院大学南京学院)
- Nanjing Institute of Information Superbahn(南京信息超级bahn研究院)
- Dobot Robotics(越疆机器人)
- School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉科学学院)
- University of Chinese Academy of Sciences(中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。