发表机构
University of Wisconsin–Milwaukee(威斯康星大学密尔沃基分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AttnFuse是一个针对注意力机制的小型可组合DSL,将RoPE等乘法前变换作为一等操作,编译为单个融合GPU内核,在RTX 3090上比flex_attention快2.10倍,在H100上接近PyTorch手工后端性能,并揭示了旋转演算规则。
AI 中文摘要
现代AI系统构建于Transformer架构之上,其核心操作——注意力机制——占据了大部分计算和内存成本。研究人员不断提出新的注意力变体以提高质量、效率或上下文长度,但每个变体目前都需要专家编写的GPU代码才能以可用速度运行。PyTorch最近的flex_attention允许研究人员用Python描述自定义注意力模式并将其编译为融合内核,但其设计仅限于在中心矩阵乘法之后应用的修改,排除了旋转位置嵌入(RoPE)——所有主流大语言模型使用的的位置编码。我们引入AttnFuse,一个用于注意力的小型DSL,使诸如RoPE之类的乘法前变换成为一等操作。研究人员组合十个高层构建块来描述一个变体,AttnFuse的编译器为整个计算生成单个融合GPU内核。在RTX 3090上,AttnFuse在RoPE+因果模式上比flex_attention实现了2.10倍的加速。在H100上,它在PyTorch手工调优后端的5%以内运行完整的Llama-3-8B训练步骤。我们的研究揭示了旋转演算:是否融合RoPE或单独应用它取决于GPU的计算与带宽比,导出的交叉点与测量结果匹配。AttnFuse证明了一个小型、注意力特定的编译器可以弥合灵活研究代码与生产内核之间的差距。
英文摘要
Modern AI systems are built on the Transformer architecture, whose core operation, attention, accounts for the majority of computation and memory cost. Researchers continually propose new attention variants to improve quality, efficiency, or context length, but each variant currently requires expert-written GPU code to run at usable speeds. PyTorch's recent flex\_attention lets researchers describe custom attention patterns in Python and compile them to fused kernels, but its design is limited to modifications applied after the central matrix multiplication, excluding Rotary Position Embedding (RoPE), the positional encoding used by every major LLM. We introduce AttnFuse, a small DSL for attention that makes pre-multiplication transformations like RoPE first-class operations. Researchers compose ten high-level building blocks to describe a variant, and AttnFuse's compiler emits a single fused GPU kernel for the entire computation. On an RTX 3090, AttnFuse achieves a 2.10$\times$ speedup over flex\_attention on the RoPE+causal pattern. On an H100, it runs a full Llama-3-8B training step within 5\% of PyTorch's hand-tuned backend. Our investigation reveals the Rotation Calculus: whether to fuse RoPE or apply it separately depends on the GPU's compute-to-bandwidth ratio, with a derived crossover that matches measurement. AttnFuse demonstrates that a small, attention-specific compiler can close the gap between flexible research code and production kernels.