arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13962cs.AR

采用高带宽ReRAM近内存架构的分散式大语言模型(LLM)服务中的MoE专家执行

MoE Expert Execution in Disaggregated LLM Serving with a High-Bandwidth ReRAM Near-Memory Architecture

Kunming Shao, Ming Zeng, Xin Yuan, Binbin Liao, Yangming Zhang, Wei Wang, Tim Kwang-Ting Cheng, Chi-Ying Tsui

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对分散式LLM服务中MoE专家执行的资源效率问题,提出ReRAM近内存架构,通过优化池化、放置与取数策略提升资源占用率,在实测中大幅降低延迟与能耗。

中文摘要 AI 辅助

注意力-前馈网络(FFN)分散化将LLM模块映射至专用资源池,为将混合专家(MoE)权重驻留在高带宽FFN资源池创造了条件。然而,解码服务水平目标(SLO)限制了运行批次规模,而稀疏路由会扩大激活专家的集合,导致权重流量分摊效果差,且路由偏差会造成冷专家资源闲置。因此,FFN资源池需在稀疏集合下提供权重读取带宽密度,并在无全局共享结构的情况下恢复资源占用率。本文提出一种ReRAM近内存架构,将专家权重驻留在高带宽本地读取路径后。该设计将实际模型利用率(MFU)分解为理想MFU和资源占用率,通过有界核心本地多播池、共激活感知放置及负载感知取数恢复资源占用率,并根据诱导需求确定各通信层级的规模。针对Qwen3.5-35B-A3B、Qwen3.5-397B-A17B和GLM-5.2的实测与建模研究显示,4路池化将资源占用率从0.328提升至0.519,在峰值计算量相同的情况下,与H20相比,每token的FFN池延迟降低9.5倍,权重移动能耗降低20倍;H20注意力+ReRAM-FFN系统的解码每token输出时间(TPOT),与同构H20池相比分别降低1.25-4.0倍、2.4-10.3倍和2.5-10.4倍。

英文摘要

Attention-FFN disaggregation maps LLM modules to specialized pools, creating an opening to keep Mixture-of-Experts (MoE) weights resident in a high-bandwidth FFN pool. Decode SLOs, however, cap the run-batch while sparse routing expands the activated-expert union, so weight traffic amortizes poorly and routing skew idles cold-expert resources. The FFN pool must therefore deliver weight-read bandwidth density under sparse unions and recover occupancy under skew without a global sharing fabric. We present a ReRAM near-memory architecture that keeps expert weights resident behind high-bandwidth local reads. The design factors actual MFU into ideal MFU and occupancy, recovers occupancy with bounded core-local multicast pooling, coactivation-aware placement, and load-aware fetch, and sizes each communication level from induced demand. A measured + modeled study on Qwen3.5-35B-A3B, Qwen3.5-397B-A17B, and GLM-5.2 shows that side-4 pooling raises occupancy from 0.328 to 0.519 and, at iso-peak compute, lowers per-token FFN-pool latency by 9.5x versus H20 with 20x lower weight-movement energy; an H20-attention + ReRAM-FFN system reduces decode TPOT by 1.25-4.0x, 2.4-10.3x, and 2.5-10.4x versus a homogeneous H20 pool.

↑