发表机构
Nanjing Vocational College of Information Technology(南京信息技术职业学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对3.5D MoE推理中专家热度偏差问题,提出HCRMap框架,基于专家热度等因素动态管理专家副本,映射令牌组,有效缓解通信、内存和队列瓶颈,实验显示其能显著降低端到端延迟。
AI 中文摘要
混合专家(MoE)大语言模型在推理过程中仅激活少量专家,但令牌路由会导致持续的专家热度偏差:一小部分热专家持续接收大多数令牌,而其余专家负载较轻。在3.5D多小芯片系统上,这种偏差不仅会导致计算不平衡,还会加剧通信、内存带宽、I/O和执行队列的压力。因此,核心问题不仅是减少令牌移动,还要在不同内存层动态放置和重用热专家副本。本文提出了HCRMap,一种用于3.5D MoE推理中压力感知专家副本管理的热专家驻留映射框架。基于专家热度、权重加载成本、迁移开销和运行时资源压力,HCRMap动态确定哪些专家应被提升、保留、降级或逐出。然后将路由的令牌组映射到合适的驻留副本,从而共同缓解通信、内存和队列瓶颈。实验结果表明,HCRMap在预填充和解码阶段分别比Hydra降低了43.6%和43.0%的端到端延迟;比MoEntwine降低了34.5%和33.1%;比PIMoE降低了46.7%和46.0%。
英文摘要
Mixture-of-Experts (MoE) large language models (LLM) activate only a small number of experts during inference, but token routing introduces persistent expert hotness skew: a small set of hot experts continuously receives most tokens, while the remaining experts are lightly loaded. On 3.5D multi-chiplet systems, this skew not only causes compute imbalance but also amplifies pressure on communication, memory bandwidth, I/O, and execution queues. Therefore, the core problem is not simply to reduce token movement, but to dynamically place and reuse hot expert replicas across different memory tiers. This paper proposes HCRMap, a hot expert residency mapping framework for pressure-aware expert replica management in 3.5D MoE inference. Based on expert hotness, weight loading cost, migration overhead, and runtime resource pressure, HCRMap dynamically determines which experts should be promoted, retained, demoted, or evicted. It then maps routed token groups to suitable resident replicas, thereby jointly mitigating communication, memory, and queue bottlenecks. Experimental results show that HCRMap reduces end-to-end latency by 43.6% and 43.0% over Hydra in the prefill and decode stages, respectively; by 34.5% and 33.1% over MoEntwine; and by 46.7% and 46.0% over PIMoE.
Comments15 pages, 8 figures, 2 tables