发表机构
Trinity College Dublin; University of Thessaly(都柏林三一学院; 色萨利大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种流水线FPGA架构,通过自定义行式存储和双路径计算策略,高效加速Longformer中的带状稀疏矩阵-稠密矩阵乘法,实现高吞吐和低功耗。
AI 中文摘要
稀疏注意力机制对于处理长输入序列的Transformer模型变得越来越重要,因为与完全自注意力相比,其计算和内存复杂度更低。Longformer通过滑动窗口注意力机制实现这一目标,该机制产生结构化的带状稀疏注意力矩阵。然而,现有的稀疏Transformer加速器主要针对注意力生成或非结构化稀疏性,使得结构化稀疏注意力的稀疏矩阵-稠密矩阵乘法(SpMM)在很大程度上未被探索。本文提出了一种用于加速Longformer中带状SpMM的流水线FPGA架构。所提出的设计利用Longformer注意力矩阵的可预测稀疏模式,通过自定义的行式存储方案和隐式索引,消除了传统稀疏矩阵格式的开销,同时实现了规则的内存访问。该架构采用并行处理单元、流水线加法器树和双路径计算策略,以最大化吞吐量和硬件利用率。该加速器使用Verilog实现,并在使用Vivado 2024.2的RFSoC平台上进行评估,在初始11个周期的延迟后,每个时钟周期维持一个完整的点积结果,同时功耗保持在2.9W以下。在100 MHz频率下运行,该设计每秒实现超过1亿个点积输出,证明了直接利用结构化稀疏性加速稀疏Transformer的有效性。
英文摘要
Sparse attention mechanisms have become increasingly important for transformer models processing long input sequences due to their lower computational and memory complexity compared to full self-attention. Longformer achieves this through a sliding-window attention mechanism that produces a structured banded sparse attention matrix. However, existing sparse transformer accelerators primarily target attention generation or unstructured sparsity, leaving sparse matrix--dense matrix multiplication (SpMM) for structured sparse attention largely unexplored. This paper presents a pipelined FPGA architecture for accelerating banded SpMM in Longformer. The proposed design exploits the predictable sparsity pattern of Longformer's attention matrix through a custom row-wise storage scheme with implicit indexing, eliminating the overhead of conventional sparse matrix formats while enabling regular memory accesses. The architecture employs parallel processing elements, pipelined adder trees, and a dual-path computation strategy to maximize throughput and hardware utilization. Implemented in Verilog and evaluated on an RFSoC platform using Vivado 2024.2, the accelerator sustains one complete dot-product result per clock cycle after an initial latency of 11 cycles while maintaining power consumption below 2.9 W. Operating at 100 MHz, the design achieves over 100 million dot-product outputs per second, demonstrating the effectiveness of directly exploiting structured sparsity for sparse transformer acceleration.
CommentsAccepted at ICECS 2026