LowRank-SSM:面向FPGA上秩约简Mamba加速的软硬件协同设计
LowRank-SSM: Hardware-Software Co-Design for Rank-Reduced Mamba Acceleration on FPGA
浏览论文内容
中文总结 AI 辅助
LowRank-SSM是面向FPGA的软硬件协同设计框架,通过软件的秩约简与硬件的双路径投影等设计,在保证精度的同时,使Mamba类SSM的FPGA部署吞吐量提升2.19倍、能效提升2.03倍。
中文摘要 AI 辅助
状态空间模型(SSM)如Mamba和Mamba-2实现了线性时间的自回归推理,使其对延迟敏感且资源受限的部署场景具有吸引力。然而,它们的大输入和输出投影层带来了二次方的权重内存和片外带宽成本,这成为FPGA实际部署的瓶颈,在序列长度为1024及以上时,该成本占每个令牌运行时间的60%以上。现有加速器通过量化或激活稀疏性降低这种开销,但都未将投影秩作为显式硬件设计变量,从而未探索系统的精度-吞吐量权衡。我们提出LowRank-SSM,一种填补该空白的软硬件协同设计框架。在软件方面,我们通过训练后截断SVD分解输入和输出投影权重,并引入贪心带秩分配算法,该算法搜索每个带的秩向量,在满足用户指定精度约束的同时最小化权重存储。在硬件方面,我们将所得的因式分解投影映射到FPGA上的全流水线加速器,该加速器具有双路径投影(低秩路径和全秩路径)、融合选择性扫描单元以及五个独立的AXI主束,可在无总线争用的情况下充分利用DDR4带宽。每个带的运行时秩掩码支持所有64层的混合秩执行,且无架构开销。在400 MHz的Xilinx Versal VC1902上,部署的混合秩INT8设计达到7.89令牌/秒,在可比功耗和精度下,相比现有最优方法(SOTA)实现了2.19倍的吞吐量提升和2.03倍的能效提升。
英文摘要
State Space Models(SSMs) such as Mamba and Mamba-2 achieve linear-time autoregressive inference, making them attractive for latency-sensitive and resource-constrained deployment. Yet their large input and output projection layers impose quadratic weight memory and off-chip bandwidth costs that bottleneck practical FPGA deployment, accounting for over 60% per-token runtime at sequence lengths of 1,024 and beyond. Existing accelerators reduce this overhead through quantization or activation sparsity, but none treat projection rank as an explicit hardware design variable, leaving a systematic accuracy-throughput trade-off unexplored. We present LowRank-SSM, a hardware-software co-design framework that closes this gap. On the software side, we decompose the input and output projection weights via post-training truncated SVD and introduce a greedy bandwise rank-allocation algorithm that searches for the per-band rank vector that minimizes weight storage while respecting a user-specified accuracy constraint. On the hardware side, we map the resulting factored projections onto a fully-pipelined accelerator on an FPGA, featuring a dual-path projection(low-rank path and full-rank path), a fused selective-scan unit, and five independent AXI master bundles that saturate DDR4 bandwidth without bus contention. A per-band runtime rank mask enables mixed-rank execution across all 64 layers with zero architectural overhead. On Xilinx Versal VC1902 at 400 MHz, the deployed mixed-rank INT8 design achieves 7.89~tokens/s, representing a ${2.19\times}$ throughput improvement and ${2.03\times}$ energy-efficiency improvement over SOTA at comparable power and accuracy.