arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种具有列向压缩的灵活稀疏感知 FPGA 加速器,用于高效 CNN 推理

A Flexible Sparsity-Aware FPGA Accelerator with Column-Wise Compression for Efficient CNN Inference

Amirhossein Zarei, Shervin Vakili

arXiv 2607.19248首次发表:更新:

AI 中文总结

研究针对资源受限平台 CNN 加速难题,提出 SparHiXcel-v2 加速器,通过列向内核压缩技术及硬件 - 算法协同设计框架,在稀疏灵活性与硬件效率间平衡,经实验在吞吐量和能效上显著提升,精度有适度下降。

AI 中文摘要

在资源受限平台上高效加速卷积神经网络(CNN)具有挑战性,因为稀疏模式不规则且硬件开销大。非结构化稀疏虽模型精度高,但硬件映射效率低;结构化稀疏虽执行简单,但灵活性降低。本文提出 SparHiXcel-v2,一种基于 FPGA 的高性价比且高度可配置的 CNN 加速器,在稀疏灵活性和硬件效率间取得更好平衡。其架构围绕可扩展二维 MAC 阵列构建,引入列向内核压缩技术,以最小硬件开销处理不规则稀疏模式。还提出硬件 - 算法协同设计框架,包括排序优化方案和针对微架构的多阶段结构化剪枝与恢复算法。对 VGG16 和 ResNet18 的广泛评估表明,通过这些优化,SparHiXcel-v2 在处理吞吐量和能源效率上有显著提升。在结构化稀疏模式下,该加速器在性价比高的 AMD Kintex UltraScale+ FPGA 上,对 VGG16 可达超 2.5 TOPS 和 210 GOP/s/W,对 ResNet18 可达超 1.1 TOPS 和 72 GOP/s/W,同时精度有适度下降。

英文摘要

Efficient acceleration of convolutional neural networks (CNNs) on resource-constrained platforms remains challenging due to the irregularity of sparsity patterns and the associated hardware overhead. While unstructured sparsity offers high model accuracy, it introduces significant inefficiencies in hardware mapping, whereas structured sparsity simplifies execution at the cost of reduced flexibility. This paper presents SparHiXcel-v2, a cost-effective and highly configurable FPGA-based CNN accelerator that achieves an improved balance between sparsity flexibility and hardware efficiency. The proposed architecture is built around a scalable two-dimensional MAC array and introduces a column-wise kernel compression technique that enables efficient handling of irregular sparsity patterns with minimal hardware overhead. To further enhance performance, we propose a hardware-algorithm co-design framework, including an ordering optimization scheme and a multi-phase structured pruning and revival algorithm tailored to the microarchitecture. Extensive evaluations on VGG16 and ResNet18 demonstrate that SparHiXcel-v2 achieves substantial improvements in processing throughput and energy efficiency through the proposed optimizations. In structured sparsity mode, the accelerator reaches over 2.5 TOPS and 210 GOP/s/W for VGG16, and over 1.1 TOPS and 72 GOP/s/W for ResNet18 on a cost-effective AMD Kintex UltraScale+ FPGA, while maintaining modest accuracy degradation.

Comments19 pages, 19 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑