arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SUSpMV:基于SUS语言编写的HBM FPGA上的高频稀疏矩阵向量乘法器

SUSpMV: A High Frequency Sparse Matrix Vector Multiplier on HBM Enabled FPGA written in SUS

Lennart Van Hirtum, David Volz, Andreas Koch, Christian Plessl

arXiv 2610.10403首次发表:更新:

发表机构

University Paderborn; TU Darmstadt(帕德博恩大学; 达姆施塔特工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SUSpMV利用SUS语言特性设计深度流水线,在Alveo U280 FPGA上实现高性能SpMV加速,峰值吞吐量达144.9 GFLOPs,较先前工作提升79%。

AI 中文摘要

SUSpMV是一个使用新兴硬件描述语言SUS编写的稀疏矩阵向量乘法(SpMV)加速器。通过利用SUS独特的延迟计数和推理机制,SUSpMV能够以极小的设计复杂度开销实现非常深的流水线。这使得在Alveo U280 FPGA上以400MHz频率利用全部32个HBM通道将矩阵数据流式传输到32个计算单元(CU)的高效实现成为可能。输入和输出向量存储在DDR内存中,这允许乘法操作的零开销链接,并通过不与计算单元共享HBM带宽来提高整体系统内存带宽。一个计算单元以宽度为1024、动态选择高度(最高达32768)的瓦片处理SpMV。每个计算单元能够每周期累加最多6个独立矩阵条目的乘法,从而组合的理论峰值计算吞吐量达到153.6 GFLOPs。矩阵存储格式旨在利用给定矩阵内的密度变化,通过动态切换一种针对较密集区域优化的表示和另一种针对较稀疏区域优化的表示,允许每个条目之间最多跳转255行。评估表明,与同一平台上的先前工作相比,几何平均改进为79%。我们实现了144.9 GFLOPs的峰值吞吐量,即理论计算吞吐量的94%,而先前工作达到98 GFLOPs。

英文摘要

SUSpMV is a Sparse Matrix Vector multiplication (SpMV) accelerator written in the upcoming HDL SUS. By leveraging SUS's unique latency counting and inference mechanism, SUSpMV could be designed with very deep pipelines yet small design complexity overhead. This enables an efficient implementation which employs all 32 HBM channels on the Alveo U280 FPGA at 400MHz for streaming matrix data into 32 Compute Units (CUs). Input and output vectors are stored in DDR memory, which allows zero-overhead chaining of multiplications and increases overall system memory bandwidth by not sharing HBM bandwidth with the CUs. A CU processes the SpMV in tiles of width 1024 and a dynamically chosen height, up to 32768. Each CU is capable of accumulating the multiplications with up to 6 separate matrix entries per cycle, resulting in the combined theoretical peak computational throughput of 153.6 GFLOPs. The matrix storage format is designed to exploit density variation within a given matrix by dynamically switching between one representation optimized for denser regions, and a second optimized for sparser regions, allowing jumps of up to 255 rows between each entry. Evaluation demonstrates a 79% geometric mean improvement over prior work on the same platform. We achieve a peak throughput of 144.9 GFLOPs, or 94% of our theoretical computational throughput, compared to 98 GFLOPs reached by prior work.

CommentsAccepted at the International Conference on Field-Programmable Technology (FPT) 2026. 9 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑