arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18656cs.LGcs.PF

面向可扩展向量架构的FlashAttention

FlashAttention for Scalable Vector Architectures

Sonia Rani Gupta, Nikela Papadopoulou, Miquel Pericàs

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出面向可扩展向量架构的FlashAttention-V,将其集成到ggml并在多模型与平台上评估,该方法在预填充、解码阶段均实现显著加速,但Q8_0量化线性层的结构性瓶颈制约了长向量可扩展性。

中文摘要 AI 辅助

在CPU上运行Transformer模型的推理正变得日益重要,尤其对于小型语言模型(SLMs)而言,向量架构正成为一种颇具前景的执行载体。注意力模块因高内存带宽需求成为主要瓶颈,FlashAttention通过融合操作来优化数据局部性、减少中间内存流量,从而缓解了这一问题。本文提出FlashAttention-V,这是一种面向可扩展向量架构的分块FlashAttention,它通过利用注意力头间的并行性、头间打包(以实现对超出头维度的向量长度的高效利用),以及提升向量寄存器利用率和内存访问局部性,实现了从短向量到极长向量的高效适配。我们将FlashAttention-V集成到ggml中,并使用gem5和Banana Pi BPI-F3在TinyLlama、Llama 3.2、Qwen2.5和Pythia-410M上对其进行评估。在Banana Pi BPI-F3上,我们确认跨注意力头的循环重排序和循环展开是有效的优化原则,其性能提升随模型规模增大而扩展,且在短上下文和解码阶段最为显著。基于模拟的分析显示,FlashAttention-V在预填充阶段的512位向量长度(VL)下,相比标量FlashAttention实现了22倍至42倍的加速,当扩展至64通道和4096位VL时,还能额外获得2倍至2.5倍的增益。在解码阶段,FlashAttention-V使用512位向量长度相比标量FlashAttention实现了8倍至11倍的加速,且由于单令牌、内存受限的执行特性,其性能对向量宽度和通道数的敏感度逐渐降低。我们进一步发现Q8_0量化线性层中存在结构性瓶颈,这些瓶颈限制了长向量执行下的算术摊销,该情况在RVV和Arm SVE上均存在,表明当前的量化格式对长向量可扩展性构成了根本性挑战。

英文摘要

Inference with transformer models on CPUs is increasingly important, especially for Small Language Models (SLMs), where vector architectures are emerging as a promising execution substrate. The attention module is a major bottleneck due to high memory bandwidth requirements; FlashAttention mitigates this by fusing operations to improve data locality and reduce intermediate memory traffic. In this paper, we present FlashAttention-V, a blocked FlashAttention for scalable vector architectures that adapts efficiently from short to very long vectors by exploiting parallelism across attention heads, inter-head packing to enable efficient utilization of vector lengths beyond the head dimension, and improving vector register utilization and memory access locality. We integrate FlashAttention-V into ggml within llama.cpp and evaluate it on TinyLlama, Llama 3.2, Qwen2.5, and Pythia-410M using gem5 and a Banana Pi BPI-F3. On the Banana Pi BPI-F3, we confirm that loop reordering and loop unrolling across attention heads are effective optimization principles, scaling performance gains with larger models and most pronounced with short contexts and during decoding. Simulation-based analysis shows that FlashAttention-V achieves 22x-42x speedup over scalar FlashAttention at 512-bit VL in prefill, with an additional 2x-2.5x gain scaling to 64 lanes and 4096-bit VL. During decode, FlashAttention-V achieves 8x-11x speedup using 512-bit vector lengths over scalar FlashAttention, with performance showing diminishing sensitivity to vector width and lane count due to single-token, memory-bound execution. We further identify structural bottlenecks in Q8_0 quantized linear layers that limit arithmetic amortization under long-vector execution, consistent across RVV and Arm SVE, indicating that current quantization formats pose a fundamental challenge to long-vector scalability.

发表机构

  • Chalmers University of Technology(查尔姆斯理工大学)
  • University of Glasgow(格拉斯哥大学)

机构由 AI 辅助整理,请以论文原文为准。

↑