arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

STEEL:用于在AMD的XDNA NPU上进行节能长序列推理的稀疏感知融合注意力

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU

Victor J. B. Jung, Gagandeep Singh, Joseph Melber, Kristof Denolf, Francesco Conti, Luca Benini

arXiv 2607.09385首次发表:更新:

发表机构

AMD Research and Advanced Development (RAD); Integrated Systems Laboratory (IIS); ETH Zürich; Department of Electrical, Electronic and Information Engineering (DEI); University of Bologna(AMD研究与先进开发; 集成系统实验室; 苏黎世联邦理工学院; 电气电子信息工程系; 博洛尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对笔记本级SoC上节能推理,提出STEEL这一针对类似XDNA的NPU的FlashAttention开源实现,通过引入预填充注意力数据流公式化及稀疏感知流水线布局,降低能耗与延迟,相比CPU和GPU有显著性能提升。

AI 中文摘要

基于大语言模型的智能体在操作系统工作流程中的日益普及,增加了笔记本级片上系统(SoC)上节能推理的重要性。虽然云卸载仍然常见,但它带来了可靠性和隐私问题,对智能体工作负载尤其成问题。因此,近期的笔记本SoC集成了针对能源效率优化的神经处理引擎(NPU);然而,由于架构多样性和显式数据移动编程模型,有效地将注意力机制映射到NPU上仍然具有挑战性。在这项工作中,我们展示了STEEL,这是第一个针对类似XDNA的NPU的FlashAttention开源实现。STEEL引入了预填充注意力的数据流公式化,能够高效利用空间并行性和片上内存。此外,STEEL通过在NPU阵列上利用稀疏感知流水线布局来解决因果掩码引起的负载不平衡问题,减少同步开销并提高利用率。我们在AMD Ryzen AI 9 HX 370 SoC上评估了STEEL,并将其性能与优化的CPU和GPU实现进行了比较。实验结果表明,STEEL相对于CPU和GPU基线分别平均降低了9.17倍和1.75倍的能耗。在XDNA 1上,STEEL的平均延迟比现有技术水平降低了9.6倍,与XDNA 2上的逐层注意力实现相比,平均加速了22.8倍。

英文摘要

The growing adoption of large language model-based agents within operating system workflows has increased the importance of energy-efficient inference on laptop-class systems-on-chip (SoCs). While cloud offloading remains common, it introduces reliability and privacy concerns that are particularly problematic for agentic workloads. Recent laptop SoCs, therefore, incorporate neural processing engines (NPUs) optimized for energy efficiency; however, effectively mapping attention mechanisms onto NPUs remains challenging due to architectural diversity and explicit data-movement programming models. In this work, we present STEEL, the first open-source implementation of FlashAttention targeting XDNA-like NPUs. STEEL introduces a dataflow formulation of prefill attention, enabling efficient exploitation of spatial parallelism and on-chip memory. Furthermore, STEEL addresses the load imbalance induced by the causal mask by leveraging a sparsity-aware pipeline placement onto the NPU array, reducing synchronization overhead and improving utilization. We evaluate STEEL on the AMD Ryzen AI 9 HX 370 SoC and compare its performance against optimized CPU and GPU implementations. Experimental results show that STEEL reduces energy consumption by an average of 9.17x and 1.75x relative to CPU and GPU baselines, respectively. On XDNA 1, STEEL achieves an average 9.6x latency reduction over the prior state of the art, and delivers a 22.8x speedup on average compared to a layer-by-layer attention implementation on XDNA 2.

CommentsAccepted at IEEE COINS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑