arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ThAME:用于大语言模型专家混合的支持3D内存的异构加速器

ThAME: 3D Memory-Enabled Heterogeneous Accelerator for LLM Mixture of Experts

Pratyush Dhingra, Pramit Kumar Pal, Janardhan Rao Doppa, Partha Pratim Pande

arXiv 2607.17074首次发表:更新:

发表机构

Washington State University(华盛顿州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对传统硬件上MoE推理的瓶颈,提出ThAME这一3D异构多芯片架构,采用特殊内存芯片及计算映射策略,设计优化的片上网络通信主干,实验显示其在加速和能源效率上比现有技术有显著提升。

AI 中文摘要

专家混合(Mixture of Experts,MoE)架构已成为扩展大语言模型(LLMs)的主导范式。然而,传统硬件上的MoE推理受到三个基本瓶颈的限制,包括获取不连续专家权重所需的大量内存带宽、输入依赖令牌路由产生的非确定性散射收集流量,以及同步专家输出聚合带来的尾部延迟依赖性。为应对这些挑战,我们提出了ThAME,一种用于MoE推理的三维(3D)异构多芯片架构。ThAME采用基于铁电场效应晶体管(FeFET)的非易失性和基于DRAM的易失性内存芯片,并采用共同设计的计算映射策略,以匹配注意力机制和专家路由的不同计算配置文件。此外,我们设计了一种专门优化的片上网络通信主干,以缓解与输入依赖的MoE流量模式组合空间中的非确定性令牌路由流量相关的瓶颈。实验结果表明,ThAME在加速方面比现有技术高出15.7倍,在能源效率方面提高了9.8倍。

英文摘要

Mixture of Experts (MoE) architectures have emerged as a dominant paradigm for scaling Large Language Models (LLMs). However, MoE inference on conventional hardware is constrained by three fundamental bottlenecks. These encompass the massive memory bandwidth required to fetch non-contiguous expert weights, the non-deterministic scatter-gather traffic generated by input-dependent token routing, and the tail-latency dependency imposed by synchronous expert output aggregation. To address these challenges, we propose ThAME, a three-dimensional (3D) heterogeneous multi-chiplet architecture for MoE inference. ThAME employs Ferroelectric Field-Effect Transistor (FeFET)-based non-volatile and DRAM-based volatile memory chiplets with a co-designed compute mapping strategy that aligns the distinct computational profiles of attention mechanisms and expert routing. Furthermore, we design a specialized Network-on-Chip communication backbone optimized to mitigate the bottlenecks associated with non-deterministic token routing traffic across the combinatorial space of input-dependent MoE traffic patterns. Experimental results demonstrate that ThAME outperforms state-of-the-art counterparts by up to 15.7x in terms of speedup and improves energy efficiency by up to 9.8x.

CommentsAccepted for Publication in IEEE/ACM Embedded Systems Week (ESWEEK-26)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑