arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10803cs.AR

针对昇腾NPU上动态形状的自适应矩阵乘法

Adaptive Matrix Multiplication for Dynamic Shapes on Ascend NPUs

Yuhang Zhou, Jiang Peng, Qianyu Jiang, Zhibin Wang, Xinghui Tian, Jianwei Zhou, Songxiang Zhu, Jingyi Zhang, Junsong Wang, Chen Tian

首次发表
浏览论文内容

中文总结 AI 辅助

针对昇腾NPU动态形状下矩阵乘法的泛化危机,提出AdaptCore自适应框架,通过解耦优化、硬件感知分块等实现高效调度,在8万种输入形状及端到端模型上取得显著加速效果。

中文摘要 AI 辅助

矩阵乘法(MatMul)面临着由高度动态张量形状引发的“泛化危机”,这一危机在昇腾NPUs上尤为突出,因为昇腾NPUs采用显式控制的架构和严格的物理约束,使得现有的以GPU为中心的优化方法失效。为解决这一问题,我们提出了AdaptCore,一个用于在昇腾NPUs上实现通用高性能矩阵乘法的自适应框架。AdaptCore将算子优化系统地解耦为空间分块和指令编排:首先将动态形状映射到硬件感知的二维分块分类体系,以平衡片上容量限制与多核并行性;此外,它将可组合的优化库与确定性分析性能模型相结合,通过数学评估硬件状态变化,主动选择并缓存最优实现,实现了O(1)开销的运行时调度。评估结果显示,AdaptCore在80000种输入形状上实现了平均1.85倍的加速,在代表性端到端模型上相较于高度调优的原生厂商库(ACLNN)实现了最高1.48倍的加速。

英文摘要

Matrix Multiplication (MatMul) faces a "generalization crisis" driven by highly dynamic tensor shapes. This crisis is particularly acute on Ascend NPUs, where explicitly controlled architectures and strict physical constraints render existing GPU-centric optimizations ineffective. To resolve this, we propose AdaptCore, an adaptive framework for universally high-performance MatMul on Ascend NPUs. AdaptCore systematically decouples operator optimization into spatial tiling and instruction orchestration. It first maps dynamic shapes into a hardware-aware 2D tiling taxonomy to balance on-chip capacity limits and multi-core parallelism. Furthermore, it integrates a composable optimization library with a deterministic analytical performance model. By mathematically evaluating hardware state mutations, AdaptCore proactively selects and caches optimal implementations, enabling O(1) overhead runtime dispatching. Evaluations demonstrate that AdaptCore delivers a remarkable 1.85x mean speedup across 80,000 input shapes, and achieves up to a 1.48x acceleration in representative end-to-end models over the highly-tuned native vendor library (ACLNN).

↑