AI 中文总结
本文提出一种基于切比雪夫多项式的脉动阵列激活单元架构,可支持多种单变量激活函数与Softmax函数,相比CORDIC等方案,在误差、面积和功耗上均有显著优化,实现了硬件资源的共享。
AI 中文摘要
近年来,神经网络加速器因比基于CPU的平台效率更高而受到广泛关注。这些加速器通常使用不同的硬件单元处理单变量激活函数(如tanh)和多变量Softmax函数,从而错失了两者间资源共享的机会。本文提出了一种新型的基于脉动阵列的激活单元架构,该架构支持多种单变量激活函数及Softmax函数。通过利用切比雪夫多项式近似,与CORDIC基线相比,我们的激活函数单元在tanh上的平均绝对误差降低了71%,同时面积减少4.6%,功耗降低5.1%;与CORDIC和分段线性近似相比,我们的Softmax近似的KL散度分别降低了44.6%和79.0%。
英文摘要
Neural Network Accelerators have gained popularity in recent years due to their greater efficiency than CPU-based platforms. Often, these accelerators utilize different hardware units for univariate activation functions, such as tanh, and the multivariate softmax, thereby missing opportunities for resource sharing between them. In this paper, we describe a novel systolic array-based activation unit architecture that supports multiple univariate activation functions as well as the softmax function. By utilizing Chebyshev polynomial approximations, our activation function unit achieves up to 71% lower mean absolute error for tanh compared to a CORDIC baseline, while using 4.6% less area and 5.1% less power. Our softmax approximation enables a 44.6% and 79.0% lower KL divergence compared to CORDIC and a piecewise-linear approximation, respectively.
CommentsAccepted for IEEE COINS 2026