arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33614cs.AI

EAT:面向高效MoE推理的专家账户追踪器

EAT: Expert Account Tracker for Efficient MoE Inference

  • Paris Elite Institute of Technology, Shanghai Jiao Tong University(上海交通大学巴黎卓越工程师学院)
  • AGI Institute, School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学与学院通用人工智能研究院)
  • Ant Group(蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

Yuexian Li, Yifei Yang, Zouying Cao, Hai Zhao

AI总结:

提出EAT方法,利用历史感知指标和自适应阈值动态选择重要专家,减少MoE推理中激活专家数量,平均降低超25%,同时保持性能,并发现高层专家更关键。

AI中文摘要:

混合专家(MoE)模型已成为扩展Transformer模型的一种革命性方法。然而,传统的MoE架构仍然存在效率低下的问题,因为大量专家被不必要地激活。现有的减少激活专家数量的方法往往忽视了每个专家的历史表现。在本文中,我们提出了EAT,一种名为专家账户追踪器(Expert Account Tracker, EAT)的新方法,该方法利用历史感知指标和自适应阈值动态选择最重要的专家,从而在有效保持模型性能的同时减少激活专家数量。实验表明,EAT在多个模型和数据集上优于现有基线Top-P方法,与原始方法相比,激活专家数量平均减少超过25%,并且与基线相比具有更好的令牌生成速度。此外,通过仅使用9K数据,被剪枝模型的性能可以通过OPD高效恢复。另外,通过消融研究,我们发现过度减少激活专家数量会显著损害模型性能,且专家的重要性在不同层间存在差异,较高层的专家通常更为关键。

英文摘要:

Mixture-of-Experts (MoE) models have emerged as a revolutionary method to scale Transformer models. However, traditional MoE architecture still suffers from inefficiency since a large number of experts are unnecessarily activated. Existing approaches for reducing the number of activated experts often overlook the historical performance of each expert. In this paper, we propose EAT, a novel method called Expert Account Tracker (EAT), which utilizes history-awareness metrics and adaptive thresholding to dynamically select the most important experts, thereby reducing the activated expert number while effectively maintaining the model performance. Experiments show that EAT outperforms the existing baseline Top-P method across multiple models and datasets, achieving over 25% an average reduction compared to the vanilla method in the number of activated experts and performing better token generation speed compared to the baseline. Furthermore, the performance of pruned models can be efficiently recovered via OPD using only 9K data. Additionally, through ablation studies, we find that excessively reducing the number of activated experts can significantly harm model performance, and the importance of experts varies across layers, with higher-level experts being generally more critical.

↑