arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

窄矩阵乘法的向量化实现用于昇腾AI推理加速

Vectorization Of Narrow Matrix Multiplication for Ascend AI Inference Acceleration

Anton Shurygin, Aleksandr Frolov

arXiv 2609.16009首次发表:更新:

AI 中文总结

针对昇腾NPU上窄矩阵乘法中Cube单元利用率低的问题,提出利用AscendC向量指令将计算卸载至Vector单元,通过重叠AIV与AIC计算,在DeepSeek-V3推理中实现平均20%性能提升。

AI 中文摘要

本研究提出并评估了一种在华为昇腾NPU上优化矩阵乘法(MatMul)的新方法,其动机源于一个关键洞察:在矩阵-向量乘法(窄矩阵乘法)过程中,Cube单元(AIC)往往未被充分利用,而Vector单元(AIV)在算子运行的大部分时间内处于空闲状态。本文中,我们引入了MatMul算法,该算法利用AscendC的向量指令,有效地将计算从Cube单元卸载到Vector单元。该算法经过测试,并应用于加速MLA DeepSeek-V3算子的推理。通过成功重叠AIV和AIC的计算,我们的优化在单token处理场景中实现了平均20%的性能提升。尽管已有相关文档和活跃的CANN社区,我们的工作填补了关于AscendC实用优化技术的文献中的一个重要空白。

英文摘要

This research proposes and evaluates a novel approach to optimizing matrix multiplication (MatMul) on Huawei Ascend NPUs, motivated by a key insight: during matrix-vector multiplication (narrow MatMul), the Cube Unit (AIC) is often underutilized, while the Vector Unit (AIV) remains idle for most of the operator runtime. In this paper, we introduce the MatMul algorithm, which uses vector instructions of AscendC to effectively offload computations from the Cube Unit to the Vector Unit. The algorithm was tested and applied to accelerating the inference of MLA DeepSeek-V3 operator. By successfully overlapping AIV and AIC computations, our optimization showed a mean performance gain of 20% for a single token processing scenario. Our work addresses a significant gap in the literature on practical optimization techniques for AscendC, despite the availability of documentation and the active CANN community.

CommentsPublished in: 2025 IEEE International Conference on Cloud Computing Technology and Science (CloudCom). Minor corrections have been made to the published abstracts

Journal ref2025 IEEE International Conference on Cloud Computing Technology and Science (CloudCom)

DOI:10.1109/CloudCom67567.2025.11331396

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑