UniMoMo:基于专家合并的大推荐模型混合专家(MoE)加速方法
UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models
浏览论文内容
中文总结 AI 辅助
UniMoMo是一种后训练压缩框架,通过功能相似性合并MoE专家并引入层自适应保护,在Amazon Beauty等三个数据集上实现推荐MoE的高效压缩与加速,且性能损失极小。
中文摘要 AI 辅助
稀疏混合专家(MoE)层通过条件计算扩展推荐容量,但训练后的检查点仍会存储和路由全部专家库。本文研究部署问题:在明确的专家预算下将该检查点转换为更小的标准MoE,且不添加压缩专用在线模块。为解决此问题,我们提出UniMoMo,一种后训练压缩框架,被表述为约束图粗化问题。UniMoMo不依赖参数距离,而是基于专家的功能相似性对其分组,使用未标记校准集测量专家对共享推荐状态的响应相似性。为防止性能下降,我们引入层自适应保护机制,基于高流量专家的路由暴露限制其合并。在Amazon Beauty、KuaiRec和TenRec三个数据集(分别含2、4、6个MoE块)上,最终4个专家的检查点获得源相对五轮平均NDCG@10比值为99.92%--102.30%,A100实测加速比为1.28×--1.63×;激进的2个专家、top-1操作点获得比值98.36%--104.24%,加速比1.47×--2.21×。这些端点结果评估了完整转换与适配工作流,表明训练后的推荐MoE可按多个服务预算导出。
英文摘要
Sparse mixture-of-experts (MoE) layers expand recommendation capacity through conditional computation, yet a trained checkpoint still stores and routes over its full expert bank. We study a deployment problem: convert that checkpoint to a smaller standard MoE under an explicit expert budget, without adding a compression-specific online module. To address this, we introduce UniMoMo, a post-training compression framework formulated as a constrained graph coarsening problem. Rather than relying on parameter distance, UniMoMo groups experts based on their functional similarity, using an unlabeled calibration set to measure how similarly experts respond to shared recommendation states. To prevent performance degradation, we introduce a layer-adaptive protection mechanism that restricts the merging of high-traffic experts based on their routing exposure. Across Amazon Beauty, KuaiRec, and TenRec with 2, 4, and 6 MoE blocks, the final four-expert checkpoints obtain source-relative five-run mean NDCG@10 ratios of 99.92%--102.30% and measured A100 speedups of 1.28$\times$--1.63$\times$. An aggressive two-expert, top-1 operating point obtains ratios of 98.36%--104.24% and speedups of 1.47$\times$--2.21$\times$. These endpoint results evaluate the complete conversion-and-adaptation workflow and show that a trained recommendation MoE can be exported at multiple serving budgets.
发表机构
- Kuaishou Technology(快手科技)
- Hohai University(河海大学)
- Wuhan University(武汉大学)
- ByteDance, Seed(字节跳动 Seed)
- Alibaba, Ant Group(阿里巴巴蚂蚁集团)
- ByteDance, Douyin(字节跳动抖音)
- The University of Hong Kong(香港大学)
- Harvard University(哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。