arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32114cs.DCcs.AI

赋能混合注意力模型在NPU上的高效运行

Empowering Hybrid Attention Models on NPUs

Yinyuan Zhang, Daliang Xu, Xiaolong Huang, Wangsong Yin, Yun Ma, Mengwei Xu, Gang Huang

首次发表
浏览论文内容

中文总结 AI 辅助

针对混合注意力模型在边缘NPU上执行效率低的问题,提出HA-NPU系统,通过三级数据流重组优化线性注意力层,实现高达35.95倍内核加速和2.03倍端到端延迟提升。

中文摘要 AI 辅助

混合注意力模型已成为大型语言模型(LLM)(如Qwen3.5和Kimi系列)的关键架构。其内存和计算效率使其在设备端推理中极具吸引力,并与边缘神经处理单元(NPU)形成了有前景的协同效应。然而,在边缘NPU上直接执行这些混合模型无法实现上述优势,通常因线性注意力(LA)层中严重的内存系统低效和架构不匹配而成为预填充阶段的瓶颈。我们提出了HA-NPU,这是首个无需修改底层算法即可在边缘NPU上实现高效混合注意力LLM推理的系统。HA-NPU通过三个层级重组LA组件的数据流来提升执行效率:(1)在核心层级,按头维度划分工作负载并融合依赖算子,消除跨核全局内存访问;(2)在算子层级,重新排序执行顺序以立即消费中间张量,大幅降低本地缓冲区压力;(3)在张量层级,采用数据流感知的布局规划,最小化矩阵与向量处理单元之间的转换开销。与竞争性基线相比,HA-NPU实现了高达35.95倍的LA内核加速和36.14倍的能耗降低,端到端请求延迟最高提升2.03倍。源代码将在该https URL上公开提供。

英文摘要

Hybrid attention models have emerged as a crucial architecture for Large Language Models (LLMs) (e.g., the Qwen3.5 and Kimi series). Their memory and computational efficiency make them highly attractive for on-device inference, forming a promising synergy with edge Neural Processing Units (NPUs). However, naive execution of these hybrid models on edge NPUs fails to deliver these benefits, often bottlenecking the prefill stage due to severe memory-system inefficiencies and architectural mismatches within the linear attention (LA) layers. We present HA-NPU, the first system to enable efficient hybrid attention LLM inference on edge NPUs without modifying the underlying algorithms. HA-NPU enhances execution efficiency by reorganizing the dataflow of the LA components across three levels: (1) At the core level, it partitions workloads by the head dimension and fuses dependent operators, eliminating cross-core global memory accesses; (2) At the operator level, it reorders execution to consume intermediate tensors immediately, drastically minimizing local-buffer pressure; (3) At the tensor level, it employs dataflow-aware layout planning to minimize transformation overhead between matrix and vector processing units. Compared to competitive baselines, HA-NPU achieves up to 35.95$\times$ LA kernel speedup and 36.14$\times$ energy reduction, delivering up to 2.03$\times$ faster end-to-end request latency. The source code will be made publicly available at https://github.com/yinyuanzhang/HA-NPU

发表机构

  • Peking University(北京大学)
  • Beijing University of Posts and Telecommunications(北京邮电大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑