多尺度能手:用于MXFP精度的通用FPGA张量模块
Jack of All Scales: A Versatile FPGA Tensor Block for MXFP Precisions
浏览论文内容
中文总结 AI 辅助
研究现代深度学习中MXFP格式在FPGA架构中支持有限的问题,提出针对性修改DSP模块内部张量模式架构的方法,实现对所有MXFP精度的原生支持,使DSP切片面积增加36%,平均吞吐量提高4.2倍。
中文摘要 AI 辅助
现代深度学习工作负载越来越依赖窄数值格式来提高效率和减少内存占用。最近标准化的微尺度浮点(MXFP)格式家族,包括MXFP8、MXFP6和MXFP4,为低精度推理提供了实用方法,但当前FPGA架构中的数字信号处理(DSP)模块对这些格式的原生支持有限。本文首先全面表征了Altera Agilex - 5 FPGA上MXFP点积实现,探索了一系列策略。结果表明张量模式虽对MXFP4(E2M1)和MXFP6(E2M3)有最高算术密度,但无法实现MXFP6(E3M2)或任何MXFP8精度。为此提出对DSP模块内部张量模式架构的针对性修改,以实现对所有MXFP精度的原生支持并保持向后兼容性。用开源ASAP7 PDK实现的Agilex - 5 DSP模块核心简化版本估计修改的面积成本,评估多种修改后的DSP模块设计在格式覆盖、算术密度和面积开销之间的权衡。首选设计点使DSP切片面积增加36%,仅占FPGA芯片总面积的1.8%。通过比较所有MXFP精度的脉动阵列矩阵乘法器实现,评估增强DSP模块对设备级的影响,结果表明在所有支持的MXFP格式上平均吞吐量提高了4.2倍。
英文摘要
Modern deep learning workloads increasingly rely on narrow numerical formats to improve efficiency and reduce memory footprint. The recently standardized microscaling floating-point (MXFP) family of formats, including MXFP8, MXFP6, and MXFP4, offers a practical approach to low-precision inference, yet the digital signal processing (DSP) blocks in current FPGA architectures offer limited native support for these formats. In this work, we first present a comprehensive characterization of MXFP dot product implementations on Altera Agilex-5 FPGAs, exploring a range of strategies spanning pure soft logic, DSP blocks in fixed-point, floating-point, and tensor modes. Our results show that while the tensor mode delivers the highest arithmetic density for MXFP4 (E2M1) and MXFP6 (E2M3), it cannot implement MXFP6 (E3M2) or any MXFP8 precisions, forcing designers to fall back to lower-density alternatives. Motivated by this gap, we propose targeted modifications to the DSP block's internal tensor-mode architecture that enable native support for all MXFP precisions while retaining backward compatibility. We estimate the area cost of these modifications using a simplified version of the Agilex-5 DSP block core implemented using the open-source ASAP7 PDK. We evaluate a variety of modified DSP block designs that present a tradeoff between format coverage, arithmetic density, and area overhead. Our preferred design point increases the DSP tile area by 36%, corresponding to only 1.8\% of the total FPGA die area. We evaluate the device-level impact of our enhanced DSP block by comparing systolic array matrix multiplier implementations across all MXFP precisions, contrasting the best-available strategies on the existing architecture against designs leveraging our modified DSP block. Our results demonstrate an average throughput improvement of 4.2x across all supported MXFP formats.