arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HDA-MoE:面向3D近内存处理的混合并行与动态自适应调度专家混合模型

HDA-MoE: Hybrid Parallelism and Dynamic, Adaptive Scheduling for Mixture-of-Experts with 3D Near-Memory Processing

Haochen Huang, Shuzhang Zhong, Shengxuan Qiu, Zhe Zhang, Shuangchen Li, Cong Li, Dimin Niu, Hongzhong Zheng, Guangyu Sun, Runsheng Wang, Meng Li

arXiv 2609.08682首次发表:更新:

发表机构

Peking University; Alibaba DAMO Academy(北京大学; 阿里巴巴达摩院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

HDA-MoE框架通过混合并行映射与动态自适应调度,优化3D近内存处理架构上的MoE推理,降低通信开销并提升计算利用率,实现显著加速。

AI 中文摘要

混合专家(MoE)架构已成为扩展大型语言模型(LLMs)的关键技术,以降低的计算成本实现高模型容量。然而,这种效率是以增加内存容量和带宽需求为代价的。最近的3D近内存处理(NMP)架构通过混合键合垂直集成内存和计算,提供高内部带宽和能效,使其成为加速MoE推理的有吸引力的选择。然而,NMP系统的分布式内存和计算组织为映射MoE工作负载带来了新的挑战。现有的并行化策略,如张量并行(TP)和专家并行(EP),要么遭受高通信成本,要么遭受不平衡的计算利用率,导致效率低下。此外,MoE模型的动态路由行为进一步使高效部署复杂化。为解决这些挑战,我们提出了HDA-MoE,一个通过混合并行部署和运行时调度优化NMP架构上MoE执行的框架。HDA-MoE集成了离线混合并行映射算法与在线动态自适应调度机制,以减少通信开销并提高计算利用率。实验结果表明,HDA-MoE相对于TP实现了1.1倍至3.4倍的加速,相对于EP实现了1.1倍至1.5倍的加速,相对于混合TP-EP计算平衡基线实现了1.1倍至3.7倍的加速,相对于HD-MoE实现了1.1倍至1.3倍的加速。源代码可在该https URL获取。

英文摘要

Mixture-of-Experts (MoE) architectures have become a key technique for scaling Large Language Models (LLMs), enabling high model capacity with reduced computational cost. However, this efficiency comes at the expense of increased memory capacity and bandwidth demands. Recent 3D Near-Memory Processing (NMP) architectures, which vertically integrate memory and compute through hybrid bonding, provide high internal bandwidth and energy efficiency, making them attractive for accelerating MoE inference. Nevertheless, the distributed memory and compute organization of NMP systems introduces new challenges for mapping MoE workloads. Existing parallelization strategies, such as Tensor Parallelism (TP) and Expert Parallelism (EP), suffer from either high communication costs or unbalanced computation utilization, leading to inferior efficiency. In addition, the dynamic routing behavior of MoE models further complicates efficient deployment. To address these challenges, we present HDA-MoE, a framework that optimizes MoE execution on NMP architectures through hybrid parallel deployment and runtime scheduling. HDA-MoE integrates an offline hybrid parallel mapping algorithm with an online dynamic and adaptive scheduling mechanism to reduce communication overhead while improving computation utilization. Experimental results show that HDA-MoE achieves a speedup of 1.1x--3.4x over TP, 1.1x--1.5x over EP, 1.1x--3.7x over the Hybrid TP-EP compute-balanced baseline, and 1.1x--1.3x over HD-MoE. Source code is available at https://github.com/PKU-SEC-Lab/HDA-MoE-TCAD26.

Comments14 pages, 24 figures, 10 tables. Accepted by IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑