arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21264cs.ARcs.LG

使用开源编译器工具为AMD XDNA NPU编程:FlashAttention案例研究

Programming AMD XDNA NPUs with Open-source Compiler Tools: A FlashAttention Case Study

Erwei Wang, Ephrem Wu, Victor J. B. Jung, Jiajie Li, Andre Rosti, Joseph Melber, Samuel Bayliss

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过FlashAttention案例,比较四种映射设计,提出基于Roofline脊点的融合策略,在AMD XDNA NPU上实现最高3.62 TFLOP/s和5.3-7.2倍能效提升。

中文摘要 AI 辅助

诸如AMD XDNA之类的空间NPU将计算瓦片放置在小容量本地存储器旁边,并将它们之间的数据移动交由软件处理。将多阶段工作负载映射到此类设备上,很大程度上取决于中间张量存放的位置。我们报告了在使用开源IRON和MLIR-AIR流程为FlashAttention做出这些选择时学到的经验。我们在XDNA 1和XDNA 2上比较了四种参考设计:一种分别运行每个算子,两种在芯片上的算子之间进行流式传输,还有一种将全部三个注意力阶段融合到单个内核中。融合内核将$\boldsymbol{QK}^{\mathsf T}$分数保存在计算瓦片本地存储器中,并通过级联互连减少部分结果,因此分数永远不会返回共享的MemTile存储器。在XDNA 2上,它在完整的端到端执行中达到3.62 TFLOP/s,是IRON设计的两倍,在2K及以上令牌数时,其能效是同一芯片上集成GPU的5.3至7.2倍。它覆盖了从BERT到DeepSeek的十二种LLM配置,最长支持128K令牌。每个存储器级别的Roofline分析解释了这一结果,并指出了何时应停止。XDNA 1的脊点较低,因此芯片上的流式传输已经达到计算受限状态:在XDNA 2上使吞吐量翻倍的相同融合在XDNA 1上几乎白费。将映射的操作强度与每个级别的脊点进行比较,可以在编写任何代码之前预测适用哪种情况。融合直到映射越过该脊点,然后停止。我们将参考设计作为维护的开源发布。

英文摘要

Spatial NPUs such as AMD XDNA place compute tiles beside small local memories and leave data movement between them to software. Mapping a multi-stage workload onto such a device is largely a question of where the intermediate tensors live. We report what we learned making those choices for FlashAttention with the open-source IRON and MLIR-AIR flows. We compare four reference designs on XDNA 1 and XDNA 2: one runs each operator separately, two stream between operators on chip, and one fuses all three attention stages into a single kernel. The fused kernel holds the $\boldsymbol{QK}^{\mathsf T}$ scores in compute-tile local memory and reduces partial results over the cascade interconnect, so the scores never return to shared MemTile memory. On XDNA 2, it reaches 3.62 TFLOP/s over complete end-to-end execution, twice the IRON design, with 5.3 to 7.2 times the energy efficiency of the integrated GPU on the same chip at 2K tokens and above. It covers twelve LLM configurations, from BERT to DeepSeek, up to 128K tokens. Roofline analysis at each memory level explains this result and shows when to stop. XDNA 1 has lower ridge points, so streaming on chip already reaches the compute-bound regime: the same fusion that doubles throughput on XDNA 2 is nearly wasted on XDNA 1. Comparing a mapping's operational intensity against each level's ridge point predicts which case applies before writing any code. Fuse until the mapping clears that ridge point, then stop. We release the reference designs as maintained open source.

发表机构

  • Advanced Micro Devices, Inc.(超威半导体公司)
  • ETH Zürich(苏黎世联邦理工学院)
  • Cornell University(康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

↑