arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向硬件高效脉动阵列设计的精度感知可变位处理单元

Precision-Aware Variable Bit Processing Elements for Hardware-Efficient Systolic Array Designs

Dantu Nandini Devi, Madhav Rao

arXiv 2608.22378首次发表:更新:

发表机构

IIIT-Bangalore(印度信息技术学院班加罗尔分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对权重静止型脉动阵列的浮点数乘法器,采用NSGA-II算法优化设计,通过部分积矩阵截断与压缩器融合,实现了脉动阵列硬件效率提升且保持相当的CNN准确率。

AI 中文摘要

脉动阵列(Systolic Arrays, SAs)已成为深度学习中矩阵运算的重要硬件加速器,而浮点数格式可在计算域间实现精度控制。本研究针对权重静止型脉动阵列(Weight Stationary Systolic Arrays)中的浮点数(Floating Point, FP)乘法器,研究近似计算技术,重点关注IEEE 754(FP32)、TensorFloat-32(TF32)及Brain Floating point(BF16)格式。通过在FP乘法器架构中融合部分积矩阵(Partial Product Matrix, PPM)列截断与正负压缩器,优化计算效率与精度间的权衡。采用NSGA-II优化算法探索庞大的设计空间,以演化FP乘法器设计,在维持可接受输出质量的同时实现显著的硬件改进。在各类应用的FP乘法器设计中均观测到显著的硬件效益,且保留了输出质量。所设计的脉动阵列中FP近似处理单元,对在MNIST、F-MNIST及CIFAR-10数据集上训练的模型,可提供相当的CNN准确率。在CIFAR-10数据集上运行模型时,与文献中对应的精确实现相比,排名前10的CNN性能的FP近似脉动阵列设计,实现了66%至92%的面积节省、60%至93%的功耗效益,延迟提升21%至54%。TF32与BF16近似脉动阵列设计也在维持相当CNN准确率的同时实现了显著增益。研究结果证实,FP乘法器设计中的定向近似可显著提升容错应用的硬件加速器效率,为当代计算架构中的硬件资源优化建立了有效方法。

英文摘要

Systolic arrays (SAs) have emerged as prominent hardware accelerators for matrix operations in deep learning, while floating point number formats enable precision control across computational domains. This research investigates approximate computing techniques for floating point (FP) multipliers in Weight Stationary Systolic Arrays, focusing on IEEE 754 (FP32), TensorFloat-32 (TF32), and Brain Floating point (BF16) formats. By integrating partial product matrix (PPM) column truncation with positive and negative compressors in the FP multiplier architecture, we optimize the trade-off between computational efficiency and accuracy. NSGA-II optimization algorithm was employed to explore the vast design space for evolving FP multiplier designs, towards achieving substantial hardware improvements while maintaining acceptable output quality. Substantial hardware benefits were observed in the FP multiplier designs across various applications, while preserving output quality. The FP approximated Processing Elements designed in the SA was found to offer comparable CNN accuracy for models trained on MNIST, F-MNIST, and CIFAR-10 dataset. The FP approximated SA designs that fall in the top 10 CNN performance offered substantial hardware gains in the range of 66% to 92% footprint savings, 60% to 93% of power benefits with 21% to 54% improvement in the delay when compared with the corresponding exact implementations mentioned in the literature for running the model trained on CIFAR-10 dataset. The TF32 and BF16 approximated SA designs also achieved substantial gains while maintaining comparable CNN accuracy. Our findings confirm that targeted approximation in FP multiplier design significantly improves the efficiency of hardware accelerators for error-tolerant applications, establishing an effective approach to hardware resource optimization in contemporary computing architectures.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑