NIFA:用于高效机器学习推理的非线性IMC增强型FPGA
NIFA: Nonlinear IMC enhanced FPGA for efficient ML inference
- Arizona State University(亚利桑那州立大学)
- Hewlett Packard Enterprise Labs(惠普企业实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对传统FPGA中IMC设计局限,提出含无ADC的IMC块的新颖FPGA架构,通过设计空间探索和高效映射,在CNN和Transformer基准测试中显著提升能源和面积效率,推进FPGA设计。
AI中文摘要:
近期的FPGA通过专用张量块和片上RAM计算提高了深度学习推理效率。基于电阻式随机存取存储器(ReRAM)的模拟内存计算(IMC)进一步提升了效率,在ReRAM交叉开关内直接执行向量矩阵乘法(VMM),相比传统数字逻辑,计算密度和能源效率提高了一个数量级。传统IMC设计仅支持静态权重VMM,将非线性操作和动态矩阵乘法(DIMM)留给FPGA结构。因此,IMC的优势主要局限于静态权重模型,而基于Transformer的模型受益有限。此外,每个IMC块内的模数转换器(ADC)消耗了超过70%的面积和功耗,进一步限制了系统效率和可扩展性。为解决这些限制,我们提出了一种新颖的FPGA架构,集成了无ADC的IMC块,用模拟内容可寻址存储器(ACAM)取代传统ADC,ACAM可在块内原生执行非线性操作。为充分利用此块,我们进行了FPGA感知的设计空间探索,确定最佳交叉开关尺寸,同时平衡FPGA面积、灵活性和DL性能,并开发了一种高效映射,利用ACAM执行DIMM操作,将IMC的适用性扩展到注意力计算。在基于CNN和Transformer的基准测试中,所提出的架构分别实现了高达40倍和1.9倍的更高能源效率,以及4.1倍和2.5倍的更高面积效率。总体而言,它显著提高了FPGA DL推理效率,并在基于Transformer的长输入序列工作负载上保持了强劲的增益增长。
英文摘要:
Recent FPGAs have improved deep learning (DL) inference efficiency through dedicated tensor blocks and in-BRAM computation. ReRAM-based analog in-memory computing (IMC) pushes efficiency further, offering an order-of-magnitude improvement in compute density and energy efficiency over conventional digital logic by performing vector-matrix multiplication (VMM) directly within the ReRAM crossbar; prior work has integrated such IMC blocks into FPGAs for DL inference. However, conventional IMC designs support only static-weight VMM, leaving nonlinear operations and dynamic matrix-matrix multiplication (DIMM) to the FPGA fabric. As a result, the benefits of IMC are largely confined to static-weight models, whereas Transformer-based models, which rely on frequent nonlinear and DIMM operations, gain only limited improvement. Moreover, the ADCs within each IMC block consume more than 70% of its area and power, further limiting system efficiency and scalability. To address these limitations, we propose a novel FPGA architecture that integrates an ADC-free IMC block, replacing the conventional ADC with analog content-addressable memories (ACAMs) that natively perform nonlinear operations inside the block. To fully exploit this block, we conduct an FPGA-aware design-space exploration that determines optimal crossbar dimensions while balancing FPGA area, flexibility, and DL performance, and we develop an efficient mapping that leverages ACAMs to carry out DIMM operations, extending the applicability of IMC to attention computation. On CNN and Transformer-based benchmarks, the proposed architecture achieves up to 40x and 1.9x higher energy efficiency and 4.1x and 2.5x higher area efficiency, respectively. Overall, it significantly improves FPGA DL inference efficiency and sustains robust gains on Transformer-based workloads across long input sequences, advancing domain-specialized FPGA design.