针对昇腾NPU上动态形状的自适应矩阵乘法
Adaptive Matrix Multiplication for Dynamic Shapes on Ascend NPUs
浏览论文内容
中文总结 AI 辅助
针对昇腾NPU动态形状下矩阵乘法的泛化危机,提出AdaptCore自适应框架,通过解耦优化、硬件感知分块等实现高效调度,在8万种输入形状及端到端模型上取得显著加速效果。
中文摘要 AI 辅助
矩阵乘法(MatMul)面临着由高度动态张量形状引发的“泛化危机”,这一危机在昇腾NPUs上尤为突出,因为昇腾NPUs采用显式控制的架构和严格的物理约束,使得现有的以GPU为中心的优化方法失效。为解决这一问题,我们提出了AdaptCore,一个用于在昇腾NPUs上实现通用高性能矩阵乘法的自适应框架。AdaptCore将算子优化系统地解耦为空间分块和指令编排:首先将动态形状映射到硬件感知的二维分块分类体系,以平衡片上容量限制与多核并行性;此外,它将可组合的优化库与确定性分析性能模型相结合,通过数学评估硬件状态变化,主动选择并缓存最优实现,实现了O(1)开销的运行时调度。评估结果显示,AdaptCore在80000种输入形状上实现了平均1.85倍的加速,在代表性端到端模型上相较于高度调优的原生厂商库(ACLNN)实现了最高1.48倍的加速。
英文摘要
Matrix Multiplication (MatMul) faces a "generalization crisis" driven by highly dynamic tensor shapes. This crisis is particularly acute on Ascend NPUs, where explicitly controlled architectures and strict physical constraints render existing GPU-centric optimizations ineffective. To resolve this, we propose AdaptCore, an adaptive framework for universally high-performance MatMul on Ascend NPUs. AdaptCore systematically decouples operator optimization into spatial tiling and instruction orchestration. It first maps dynamic shapes into a hardware-aware 2D tiling taxonomy to balance on-chip capacity limits and multi-core parallelism. Furthermore, it integrates a composable optimization library with a deterministic analytical performance model. By mathematically evaluating hardware state mutations, AdaptCore proactively selects and caches optimal implementations, enabling O(1) overhead runtime dispatching. Evaluations demonstrate that AdaptCore delivers a remarkable 1.85x mean speedup across 80,000 input shapes, and achieves up to a 1.48x acceleration in representative end-to-end models over the highly-tuned native vendor library (ACLNN).