发表机构
School of Computer Science, Beihang University; School of Integrated Circuits and Systems, Beihang University(北京航空航天大学计算机学院; 北京航空航天大学集成电路与系统学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SPIMOE是首个利用混合稀疏性协同设计异构PIM架构的MoE推理框架,结合自适应路由、块稀疏注意力和KV驱逐,实现8.35倍端到端加速并保持准确性。
AI 中文摘要
长推理混合专家(MoE)模型暴露了两个耦合的推理瓶颈:不断增长的KV缓存将关键路径转移到注意力机制,而稀疏的专家激活导致负载不均衡和硬件利用率低下。尽管存内处理(PIM)为缓解数据移动开销提供了一种有前景的方法,但现有的基于PIM的加速器通常单独优化注意力或前馈网络(FFN)。我们提出了SPIMOE,这是第一个利用混合稀疏性在异构PIM架构上进行高效MoE推理的协同设计框架。SPIMOE结合了自适应专家路由与块稀疏注意力和物理KV缓存驱逐,并将注意力和FFN在SRAM-PIM和HBM-PIM上进行分离。静态专家映射和动态子批调度进一步平衡通道负载并重叠两条路径。评估表明,与NVIDIA A100 GPU相比,SPIMOE实现了高达8.35倍的端到端加速,在MoE FFN执行上比PIMoE实现了3.33倍的加速,同时保持了与全注意力基线相当的推理准确性。
英文摘要
Long-reasoning Mixture-of-Experts (MoE) models expose two coupled inference bottlenecks: growing KV caches shift the critical path toward attention, while sparse expert activation causes load imbalance and low hardware utilization. Although Processing-in-Memory (PIM) offers a promising way to mitigate data movement overhead, existing PIM-based accelerators typically optimize attention or FFNs in isolation. We propose SPIMOE, the first co-design framework that exploits hybrid sparsity for efficient MoE inference on heterogeneous PIM architectures. SPIMOE combines adaptive expert routing with block-sparse attention and physical KV-cache eviction, and disaggregates attention and FFNs across SRAM-PIM and HBM-PIM. Static expert mapping and dynamic sub-batch scheduling further balance channel loads and overlap the two paths. Evaluations show that SPIMOE achieves up to $8.35\times$ end-to-end speedup over an NVIDIA A100 GPU and $3.33\times$ speedup in MoE FFN execution over PIMoE, while preserving reasoning accuracy comparable to full-attention baselines.