arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于暂存式张量加速器的MoE解码多引擎数据流

A Multi-Engine Dataflow for MoE Decoding on AWS Tranium

Bin Ma, Wenjie Fan, Jialin Liu, Dong Li

arXiv 2609.21137首次发表:更新:

发表机构

University of California, Merced; Yotta Labs(加州大学默塞德分校; Yotta Labs)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对暂存式张量加速器上MoE解码因专家权重移动导致的性能瓶颈,提出CARDAN方法,通过向量量化加低秩分解表示和多引擎数据流重叠DMA与计算,在五个MoE模型上提升解码速度1.15-1.7倍。

AI 中文摘要

在基于暂存式张量加速器(STA)上,混合专家(MoE)解码的性能主要受限于专家权重的移动,而计算引擎则处于空闲状态。这种数据传输难以隐藏,因为专家仅在路由之后才被确定,且在不损失质量或增加关键路径工作的情况下难以缩减。我们提出了CARDAN,它将每个专家权重矩阵表示为向量量化分量加上共享基低秩分量,并将此表示与多引擎解码数据流协同设计。该表示将专家共有与专家私有工作分离,使数据流能够在多个引擎上重叠DMA与计算。在AWS Trainium3上的五个MoE家族中,CARDAN在所有五个模型上匹配或改善了BF16教师困惑度,并将批大小为1的解码速度提升了1.15-1.31倍,相对于AWS密集MoE巨型内核,在批大小为16时提升至1.7倍。

英文摘要

Mixture-of-Experts (MoE) decoding on scratchpad-based tensor accelerators (STA) is dominated by moving expert weights while the compute engines sit idle. This traffic is hard to hide, because the experts are known only after routing, and hard to shrink without losing quality or adding critical-path work. We present CARDAN, which represents each expert-weight matrix as a vector-quantized component plus a shared-basis low-rank component and co-designs this representation with a multi-engine decoding dataflow. The representation separates expert-common from expert-private work, so the dataflow overlaps DMA with computation on several engines. Across five MoE families on AWS Trainium3, CARDAN matches or improves BF16-teacher perplexity across all five models and speeds up batch-one decoding by 1.15-1.31x over AWS dense MoE megakernels, rising to 1.7x at batch size 16.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑